This is Estuary's fork of gregrahn/tpcds-kit.
It exists to drive the source-tpc-ds capture connector in
estuary/connectors, which runs dsdgen
as a subprocess in stdout mode. The generator source remains subject to the TPC
legal notice in EULA.txt and at the top of every source file; the
fork redistributes it with that notice intact, as DuckDB and Trino do.
Changes on top of upstream, one commit each (diff master against upstream's
master to see them all):
print: the stdout option was registered as_FILTERbut checked asFILTER, so stdout mode never engaged.print: after selecting stdout the handle was overwritten by the table's NULL file pointer ("Failed to open output file"). Also flush rather than fclose stdout so a parent and its child table can both close it.parallel:-PARALLEL/-CHILDonly chunked tables of 1M rows or more and silently emitted nothing for other children of smaller tables. An explicit-PARALLELnow chunks every table. Chunked output concatenates byte-identically to unchunked output.driver: hidden-_ROWCOUNT Yflag prints the row count of-TABLEat-SCALEand exits, following the existing hidden-flag convention.build: prototypes for three K&R-style definitions so the tree compiles under gcc 14 and current clang without-std=gnu89.scale: fractional scale factors below 1, ported from DuckDB's tpcds extension. Every table starts from its 1GB row count and is multiplied by the fraction (floor of one row), except the fixed-size tables (date_dim, time_dim, catalog_page, ship_mode, income_band and similar). Row counts at sf=0.01 matchdsdgen(sf=0.01)in DuckDB for every directly generated table. Row content follows upstream dsdgen; DuckDB's embedded copy diverges from upstream in text and null generation, so DuckDB's shipped answer sets do not describe this output.params: string parameters up to 4095 characters (paths were strcpy'd into 80-byte buffers).driver: a command line over 200 characters produced a NULL dereference inReportError; it is now a warning and the recorded string is truncated.parallel: when a chunk's first row fell exactly on a day boundary,skipDaysplaced it on the following day and shifted that day's rows, so chunked output differed from a serial run. Rare at 1GB and above, frequent below. Chunks now concatenate byte-identically at every scale.
Dockerfile builds dsdgen as a static Linux binary and ships
it with tpcds.idx in an empty image, published on every push to master as
ghcr.io/estuary/dsdgen:<7-char commit sha> and :latest for linux/amd64 and
linux/arm64. Consume it with COPY --from:
COPY --from=ghcr.io/estuary/dsdgen:<sha> /dsdgen /tpcds.idx /usr/local/bin/Stream a table to stdout, always passing the distributions file explicitly:
dsdgen -SCALE 0.01 -TABLE store_sales -PARALLEL 4 -CHILD 1 -_FILTER Y -DISTRIBUTIONS /usr/local/bin/tpcds.idx
The three returns tables are emitted by their sales parent's process, interleaved on stdout; tell them apart by field count.
The official TPC-DS tools can be found at tpc.org.
This version is based on v2.10.0 and has been modified to:
- Allow compilation under macOS (commit 2ec45c5)
- Address obvious query template bugs like
- Rename
s_web_returnscolumnwret_web_site_idtowret_web_page_idto match specification. See #22 & #42.
To see all modifications, diff the files in the master branch to the version branch. Eg: master vs v2.10.0.
Make sure the required development tools are installed:
Ubuntu:
sudo apt-get install gcc make flex bison byacc git
CentOS/RHEL:
sudo yum install gcc make flex bison byacc git
Then run the following commands to clone the repo and build the tools:
git clone https://github.com/gregrahn/tpcds-kit.git
cd tpcds-kit/tools
make OS=LINUX
Make sure the required development tools are installed:
xcode-select --install
Then run the following commands to clone the repo and build the tools:
git clone https://github.com/gregrahn/tpcds-kit.git
cd tpcds-kit/tools
make OS=MACOS
Data generation is done via dsdgen. See dsdgen -help for all options. If you do not run dsdgen from the tools/ directory then you will need to use the option -DISTRIBUTIONS /.../tpcds-kit/tools/tpcds.idx. The output directory (specified via the -DIR option) must exist prior to running dsdgen.
Query generation is done via dsqgen. See dsqgen -help for all options.
The following command can be used to generate all 99 queries in numerical order (-QUALIFY) for the 10TB scale factor (-SCALE) using the Netezza dialect template (-DIALECT) with the output going to /tmp/query_0.sql (-OUTPUT_DIR).
dsqgen \
-DIRECTORY ../query_templates \
-INPUT ../query_templates/templates.lst \
-VERBOSE Y \
-QUALIFY Y \
-SCALE 10000 \
-DIALECT netezza \
-OUTPUT_DIR /tmp