datascope
Git diff for datasets. A debugger for machine-learning data and models — local-first, deterministic, offline.
Datascope is a command-line + Python developer tool that treats tabular datasets the way Git treats source files: it diffs changes, reviews quality, and debugs why your ML model behaves the way it does.
Everything runs on your machine. No data is uploaded, nothing is sent
anywhere, and every command is deterministic by default (fixed seed 42),
so two runs on the same inputs produce byte-identical output. Great for CI and
for reproducible notebooks.
- Inspect — profile a CSV/Parquet/JSON/JSONL into a health score (0–100) and per-column statistics.
- Diff — compare two snapshots of the same dataset with statistical tests.
- Debug — run ~15 ML-focused detectors over a dataset (or train/test pair) that flag leakage, contamination, label noise, ID columns, imbalance, outliers, missingness, drift, subgroups and low-value features.
- Drift — measure feature-distribution shift between a reference and a current snapshot (PSI / KS / JS-divergence / chi²).
- Check / CI — gate pipelines and exit non-zero when findings exceed a
severity threshold; ships with a ready-made GitHub Action and
fail-on. - Report — terminal, JSON, Markdown or self-contained HTML.
| Area | What it gives you |
|---|---|
| Profiling | Row/column counts, memory, per-column dtype plus inferred logical type, missing %, cardinality, distribution previews, dupes. |
| Health score | A transparent, deterministic 0–100 score from documented rule tables (see docs/health.md). |
| Dataset diff | Row-level adds/removes, schema changes, and statistical comparisons (KS, PSI, chi²) of changed columns. |
| ML debugger | Detectors for leakage, train/test contamination, label noise, spurious/id-like features, temporal leakage, class imbalance, outliers, missingness, feature drift and subgroup performance. |
| Honest model metrics | Out-of-fold cross-validation so every number estimates generalisation, not training error. |
| Pluggable detectors | Register your own detector in a few lines (datascope.detectors.register_detector). |
| Formats | CSV, TSV, Parquet, JSON, JSONL + in-memory pandas/polars DataFrame. |
| Output | terminal, json, markdown, html. JSON is machine-parseable; HTML is self-contained and offline. |
| Behaviour | Local-first, deterministic (default seed 42), no telemetry. --seed overrides. |
Requires Python 3.11+.
# from a checkout
pip install -e .
# or with the ML/model-debugging extras (scikit-learn)
pip install -e ".[ml]"
# everything, incl. plotly for optional interactive charts
pip install -e ".[all]"Core read-only profiling and diffing have no heavy ML dependencies: you get
it with a plain pip install ..
CLI sandbox after install:
datascope --help
datascope --version # datascope 1.0.0datascope inspect ./data/churn_train.csv --format terminalDataset: churn_train.csv (600 rows × 8 columns)
Health: 89 / 100
…per-column profile + issues…
Write a JSON file for scripting:
datascope inspect ./data/churn_train.csv -t churned --format json --output out.jsondatascope diff ./data/train_v1.csv ./data/train_v2.csv --target churnedShows how the support/spend/spend_6m distributions changed, marked with the
severity of each statistical signal (datascope uses PSI + KS for numerics,
JS-divergence/PSI/chi² for categories).
datascope debug ./data/churn_train.csv --target churned# PATH is the snapshot you debug; --train/--test (optional) enable
# train/test contamination and temporal-leakage checks against them
datascope debug data/current.csv --train data/train.csv --test data/test.csv --target churneddatascope drift ./data/train_v1.csv ./data/train_v2.csvdatascope check ./data/churn_train.csv --target churned --fail-on medium
echo $? # 0 => clean; 2 => worst finding >= "medium"Pair with the bundled GitHub Action (.github/workflows/datascope.yml) that
calls the same check semantics and fails the workflow when the gate trips.
Everything the CLI does is a thin wrapper over a small Python library.
from datascope import Dataset, DatasetDiff, MLDebugger, profile_dataset
# 1. Profile & health
resp = profile_dataset("churn.csv", target="churned")
print(resp.score, resp.profile.rows, resp.profile.columns_profiles[0].name)
# 2. Compare two snapshots
d = DatasetDiff("train_v1.csv", "train_v2.csv", target="churned").run()
print(d.overall, len(d.schema), len(d.numeric))
# 3. Model/data debugging
debug = MLDebugger("churn.csv", target="churned").run()
for f in debug.findings:
print(f.severity.value, f.kind.value, f.title, "->", f.message)Dataset, DatasetDiff, MLDebugger and profile_dataset accept a path, a
CSV/Parquet/JSON/JSONL file, or an in-memory pandas.DataFrame.
Add a .datascope.yaml in the working directory (or pass --config) to tune
detectors, thresholds, CI gates and the default seed:
# .datascope.yaml
target: churned
seed: 7
diff:
psi_threshold: 0.3
ks_pvalue_threshold: 0.01
outliers:
method: iqr # iqr | robust_zscore | isolation_forest
threshold: 4.0
class_balance:
imbalance_critical: 30.0
model:
cv_folds: 5
cv_max_rows: 50000
ablation:
enabled: true # opt-in: refits the model repeatedly
ci:
fail_on: medium
report_path: datascope-report.mdSee docs/configuration.md for every option and its default.
| Format | Usage |
|---|---|
terminal |
Human colour report (default). |
json |
Machine-readable result object (valid JSON on stdout). |
markdown |
Great for PR comments / Jupyter / GitBook. |
html |
Self-contained, offline single-file report (style inlined; no CDN). |
Choose with --format/-f or set the output extension (-o out.md → Markdown).
The html report is a single dependency-free document (no CDN, no external
fonts) — ideal to attach to CI or email. If you install the optional plotly
extra (.[all]), the small datascope.visualization module can also export
interactive, self-contained charts for an inspect or drift report from Python;
see docs/visualization.md.
- Default seed is
42for every command that samples or shuffles. Give--seedto override. Re-running produces identical bytes. - All computation is local. Datascope never uploads your data or sends telemetry.
- JSON/HTML reporters write plain text/files — no external network calls.
datascope/
├── src/datascope/
│ ├── cli.py # typer CLI
│ ├── config.py # .datascope.yaml parsing/validation
│ ├── dataset.py # Dataset facade (read + sample + sniff)
│ ├── models.py # core enums/models (Severity, Finding, …)
│ ├── profiling/ # single-dataset profiling + health scoring
│ ├── diff/ # old-vs-new comparison engine
│ ├── detectors/ # pluggable ML debugger detectors + registry
│ ├── debugging/ # MLDebugger orchestrator / report assembly
│ ├── reporting/ # terminal/json/markdown/html renderers
│ ├── metrics/ # statistical tests (KS/PSI/JS/chi²/AUC/…)
│ ├── loaders/ # CSV/Parquet/JSON/JSONL readers
│ └── utils/ # typing/serialization helpers
├── tests/ # pytest suite (targets ≥85% branch coverage)
├── docs/ # configuration, health, architecture, detectors
├── .github/workflows/ # CI that lints + tests + runs datascope check
└── pyproject.toml
Contributions are welcome — see CONTRIBUTING (and the security notes in SECURITY). In brief:
pip install -e ".[dev]"
ruff check src && ruff format --check src
mypy src/datascope
pytest --cov=datascope --cov-branchThe detector Protocol (src/datascope/detectors/base.py) plus the registry
make it easy to add a new detector without touching the orchestrator.
docs/configuration.md— every option and default.docs/health.md— the transparent health-scoring rules.docs/architecture.md— how the pieces fit together.docs/detectors.md— what each built-in detector does and when it fires.docs/visualization.md— the optional interactive Plotly charts API.
Apache-2.0. This is an open-source project; contributions are subject to the terms of the license.

