Download TabbyDB. Free Trial.
Trial license valid for 30 days. No code changes required. Full rollback in 30 seconds.
Need a Trial License?
TabbyDB requires a license file to run. Request yours — license.json will be emailed to you instantly.
Add these to your spark-defaults.conf
spark.tabbydb.license.mode offline spark.tabbydb.license.file </path/to/your/license.json>
Download TabbyDB as a complete Spark installation. 100% compatible with the corresponding Apache Spark version.
Apache Iceberg Runtime — Drop-in Replacement
The iceberg-tabbydb-runtime jar unlocks Broadcast Hash Join key pushdown at the scan level when querying Iceberg tables — unavailable in the standard iceberg-spark-runtime. On a single-node M4 Mac, 50GB non-partitioned Iceberg testing shows 75%+ improvement (3101.7s → 695.6s) for queries sorted on the date column. Larger benchmarks at 1 TB and 2 TB are underway.
- •Already included in the Fresh Install (.tgz) of TabbyDB 4.1.3 as
$SPARK_HOME/jars/iceberg-tabbydb-runtime-4.1.3_2.13-1.11.0.jar— no separate download needed. - •For the Convert Existing Spark path: download the jar below and replace your existing
iceberg-spark-runtime-4.1_2.13-*.jarwith it. - •The default
iceberg-spark-runtimejar remains 100% compatible with TabbyDB 4.1.3 — you just won't get the Broadcast Hash Join scan-level pushdown gains without this drop-in replacement.
Run TPC-DS Benchmarks Locally
Everything needed to generate TPC-DS data, create tables, and run benchmarks on both stock Spark 4.1.3 and TabbyDB on a single machine. See the FAQ walkthrough for step-by-step instructions.
Apply this if make fails while building the databricks/tpcds-kit tool. Run the following from the tpcds-kit checkout directory (not the tools/ subdirectory):
patch -p1 < tpcds-toolkit-patch.txttpcds-toolkit-patch.txt ↓
Pre-built for Spark 4.1.3 with non-partitioned, date-sorted table support built in. Use the jar directly, or build from the databricks/spark-sql-perf source by applying the patch and running the build from the checkout directory:
patch -p1 < non-partitioned-tables.patch sbt clean package -J-Djava.security.manager=allow
The built jar will be at target/scala-2.13/spark-sql-perf_2.13-0.5.1-SNAPSHOT.jar. Pass this path to --jars when launching spark-shell.
Tuned spark-defaults.conf and spark-env.sh for a single-node cluster. Drop into $SPARK_HOME/conf/ and adjust executor counts to match your hardware.
Copy-paste directly into spark-shell. Set the path variables at the top of each script before running.
- •Now tracks the Apache Spark 4.1.3 release
- •Once-only application of expensive optimizer rules to complex repeated sub-expressions regardless of batch iterations (SPARK-36786)
- •Improved Broadcast Hash Join key pushdown performance at scan level (SPARK-44662)
- •New iceberg-tabbydb-runtime-4.1.3 jar — a drop-in replacement for the standard
iceberg-spark-runtimethat extends Broadcast Hash Join key pushdown to Iceberg table scans, enabling deep file-level pruning before data is read. Download it from the Iceberg Runtime section below. - •Dynamic pruning at scan level using Broadcast Hash Join keys is now also supported for cached in-memory tables, in addition to the existing support for Hive and Iceberg Parquet-formatted tables.
- •Support for auto-caching/uncaching of common subplans within a query to speed up execution. Off by default — enable with
spark.sql.analyzer.commonPlanAutoCachingEnabled=true. This uses more memory and should be treated as experimental for now.
Compare Performance in Real Time
Run the same query on stock Spark vs TabbyDB in our hosted Zeppelin notebooks.