Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

TensorFIM — Reproducibility Package

This package reproduces the experiments of:

TensorFIM: Exact Maximal Frequent Itemset Mining on Tensor Cores (TKDE submission)

The central claim is verifiable bit-exactly on any NVIDIA GPU from sm_80 upward: tensor-core (b1 BMMA) support counting produces identical maximal frequent itemsets to the classical bitmap engine, and every script in this package re-checks that identity against archived gold references.

Environment

  • GPU: any NVIDIA GPU with compute capability >= 8.0 (Ampere or newer). Reference numbers were measured on an RTX 3060 Ti (sm_86, 38 SMs).
  • CUDA toolkit >= 12.0, a C++17 host compiler (MSVC on Windows, GCC on Linux).
  • Python >= 3.8 for the driver scripts (standard library only).
  • Prebuilt Windows x64 binaries (code/engine/run/) are included for convenience; they are statically linked and run without a CUDA installation.

Layout

code/
  microbench/   bmma_bench.cu — isolated BMMA vs. bitmap-loop microbenchmark (Table I, Fig. 2)
  batch_proto/  batch_proto.cu — wave-batching prototype on synthetic data (Section V-C)
  engine/       kernel_bitmap.cu (baseline) + kernel_bmma.cu (TensorFIM), build scripts
scripts/
  run_all_configs.py       14 dataset-threshold configurations, both engines (Tables IV/V, Figs. 3/5)
  download_chainstore.py   fetch chain-store from the public SPMF dataset library
  download_webdocs.py      fetch webdocs from the release-asset mirror (FIMI original offline)
  run_ablation.py          wave-size sweep + split-k on/off (Table VII)
  run_threshold_sweep.py   accidents/kosarak/connect across thresholds (Fig. 7)
  verify_all.py            set-identity check of all produced results vs. gold
data/
  small/          the seven small-to-mid FIMI benchmarks
  gold_results/   reference MFI sets for every configuration (the correctness anchor)
  make_replica.py, strip_tcga.py, DOWNLOAD.md   large-dataset generation/provenance
portable/       self-contained cross-platform validation bundle (Section V-K):
                auto-detects the GPU, builds from source or JITs the shipped PTX,
                runs the reference configurations, writes back MD5-manifested results.
                Double-click RUN_ME.bat on any Windows host with an sm_80+ GPU;
                datasets are drawn from data/ or generated by data/make_replica.py.

Quickstart

# 1. build the engines (or use the prebuilt binaries in code/engine/run/)
cd code/engine && ./build.bat          # nvcc -O3 -arch=native; MSVC + CUDA on Windows

# 2. microbenchmark: BMMA counting throughput (Table I, Fig. 2)
cd ../microbench && ./build.bat && ./bmma_bench.exe

# 3. end-to-end: all 14 configurations, 5 interleaved reps, set-verified
#    (chain-store and webdocs first: python scripts/download_chainstore.py / download_webdocs.py)
cd ../.. && python scripts/run_all_configs.py

# 4. ablation and threshold sweep
python scripts/run_ablation.py
python scripts/run_threshold_sweep.py

# 5. re-verify every result against the gold references
python scripts/verify_all.py

Expected outcome of every correctness gate: ALL SETS IDENTICAL. Timings vary with hardware; the reference medians (RTX 3060 Ti) are below.

Claim-to-command map

Paper claim Command Reference (RTX 3060 Ti)
BMMA sustains 56.2 Tbit-op/s, 25.8x geomean over bitmap loop code/microbench/bmma_bench.exe Table I, Fig. 2
Prototype end-to-end 31.0x, identical output code/batch_proto/batch_proto.exe Section V-C, Table II
End-to-end wins on 13 of 14 configurations, up to 12.4x vs the bitmap baseline scripts/run_all_configs.py Tables IV/V, Figs. 3/5
Bit-exact output on all configurations every script + verify_all.py Table III
Phase breakdown: BMMA dominates post-compaction engine logs (BMMA phase ms: line) Fig. 4, Section V-F
Amdahl ceiling 25.2x on pumsb_x256 cost model from phase logs Section V-F
Split-k off = 24.9x slower on pumsb_x256 scripts/run_ablation.py Table VII
Threshold-sensitivity crossover scripts/run_threshold_sweep.py Fig. 7

Notes

  • The engines exit with code 1 after a successful run (historical convention); success is the produced <dataset>-<thr>=Results.txt, not the exit code.
  • pumsb_x256 and webdocs are large (4.3 GB / 1.5 GB); see data/DOWNLOAD.md. All other configurations run in seconds.
  • Baselines: GMiner (https://github.com/opensourcesavvy/GMiner; see paper ref [8]) and FPmax* (paper ref [44]) are third-party codes and not redistributed here; Table V's baseline columns were measured on the same host with the protocols described in Section V-A.

About

Frequent itemset mining on GPU Tensor Cores via bitmap-WMMA hybridization (TKDE reproduction package)

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages