This package reproduces the experiments of:
TensorFIM: Exact Maximal Frequent Itemset Mining on Tensor Cores (TKDE submission)
The central claim is verifiable bit-exactly on any NVIDIA GPU from sm_80 upward: tensor-core (b1 BMMA) support counting produces identical maximal frequent itemsets to the classical bitmap engine, and every script in this package re-checks that identity against archived gold references.
- GPU: any NVIDIA GPU with compute capability >= 8.0 (Ampere or newer). Reference numbers were measured on an RTX 3060 Ti (sm_86, 38 SMs).
- CUDA toolkit >= 12.0, a C++17 host compiler (MSVC on Windows, GCC on Linux).
- Python >= 3.8 for the driver scripts (standard library only).
- Prebuilt Windows x64 binaries (
code/engine/run/) are included for convenience; they are statically linked and run without a CUDA installation.
code/
microbench/ bmma_bench.cu — isolated BMMA vs. bitmap-loop microbenchmark (Table I, Fig. 2)
batch_proto/ batch_proto.cu — wave-batching prototype on synthetic data (Section V-C)
engine/ kernel_bitmap.cu (baseline) + kernel_bmma.cu (TensorFIM), build scripts
scripts/
run_all_configs.py 14 dataset-threshold configurations, both engines (Tables IV/V, Figs. 3/5)
download_chainstore.py fetch chain-store from the public SPMF dataset library
download_webdocs.py fetch webdocs from the release-asset mirror (FIMI original offline)
run_ablation.py wave-size sweep + split-k on/off (Table VII)
run_threshold_sweep.py accidents/kosarak/connect across thresholds (Fig. 7)
verify_all.py set-identity check of all produced results vs. gold
data/
small/ the seven small-to-mid FIMI benchmarks
gold_results/ reference MFI sets for every configuration (the correctness anchor)
make_replica.py, strip_tcga.py, DOWNLOAD.md large-dataset generation/provenance
portable/ self-contained cross-platform validation bundle (Section V-K):
auto-detects the GPU, builds from source or JITs the shipped PTX,
runs the reference configurations, writes back MD5-manifested results.
Double-click RUN_ME.bat on any Windows host with an sm_80+ GPU;
datasets are drawn from data/ or generated by data/make_replica.py.
# 1. build the engines (or use the prebuilt binaries in code/engine/run/)
cd code/engine && ./build.bat # nvcc -O3 -arch=native; MSVC + CUDA on Windows
# 2. microbenchmark: BMMA counting throughput (Table I, Fig. 2)
cd ../microbench && ./build.bat && ./bmma_bench.exe
# 3. end-to-end: all 14 configurations, 5 interleaved reps, set-verified
# (chain-store and webdocs first: python scripts/download_chainstore.py / download_webdocs.py)
cd ../.. && python scripts/run_all_configs.py
# 4. ablation and threshold sweep
python scripts/run_ablation.py
python scripts/run_threshold_sweep.py
# 5. re-verify every result against the gold references
python scripts/verify_all.pyExpected outcome of every correctness gate: ALL SETS IDENTICAL.
Timings vary with hardware; the reference medians (RTX 3060 Ti) are below.
| Paper claim | Command | Reference (RTX 3060 Ti) |
|---|---|---|
| BMMA sustains 56.2 Tbit-op/s, 25.8x geomean over bitmap loop | code/microbench/bmma_bench.exe |
Table I, Fig. 2 |
| Prototype end-to-end 31.0x, identical output | code/batch_proto/batch_proto.exe |
Section V-C, Table II |
| End-to-end wins on 13 of 14 configurations, up to 12.4x vs the bitmap baseline | scripts/run_all_configs.py |
Tables IV/V, Figs. 3/5 |
| Bit-exact output on all configurations | every script + verify_all.py |
Table III |
| Phase breakdown: BMMA dominates post-compaction | engine logs (BMMA phase ms: line) |
Fig. 4, Section V-F |
| Amdahl ceiling 25.2x on pumsb_x256 | cost model from phase logs | Section V-F |
| Split-k off = 24.9x slower on pumsb_x256 | scripts/run_ablation.py |
Table VII |
| Threshold-sensitivity crossover | scripts/run_threshold_sweep.py |
Fig. 7 |
- The engines exit with code 1 after a successful run (historical convention);
success is the produced
<dataset>-<thr>=Results.txt, not the exit code. - pumsb_x256 and webdocs are large (4.3 GB / 1.5 GB); see
data/DOWNLOAD.md. All other configurations run in seconds. - Baselines: GMiner (https://github.com/opensourcesavvy/GMiner; see paper ref [8]) and FPmax* (paper ref [44]) are third-party codes and not redistributed here; Table V's baseline columns were measured on the same host with the protocols described in Section V-A.