Skip to content

Repository files navigation

datascope

DataScope help

Git diff for datasets. A debugger for machine-learning data and models — local-first, deterministic, offline.

License: Apache-2.0 Python 3.11+ status

Datascope is a command-line + Python developer tool that treats tabular datasets the way Git treats source files: it diffs changes, reviews quality, and debugs why your ML model behaves the way it does.

DataScope example

Everything runs on your machine. No data is uploaded, nothing is sent anywhere, and every command is deterministic by default (fixed seed 42), so two runs on the same inputs produce byte-identical output. Great for CI and for reproducible notebooks.

  • Inspect — profile a CSV/Parquet/JSON/JSONL into a health score (0–100) and per-column statistics.
  • Diff — compare two snapshots of the same dataset with statistical tests.
  • Debug — run ~15 ML-focused detectors over a dataset (or train/test pair) that flag leakage, contamination, label noise, ID columns, imbalance, outliers, missingness, drift, subgroups and low-value features.
  • Drift — measure feature-distribution shift between a reference and a current snapshot (PSI / KS / JS-divergence / chi²).
  • Check / CI — gate pipelines and exit non-zero when findings exceed a severity threshold; ships with a ready-made GitHub Action and fail-on.
  • Report — terminal, JSON, Markdown or self-contained HTML.

Features

Area What it gives you
Profiling Row/column counts, memory, per-column dtype plus inferred logical type, missing %, cardinality, distribution previews, dupes.
Health score A transparent, deterministic 0–100 score from documented rule tables (see docs/health.md).
Dataset diff Row-level adds/removes, schema changes, and statistical comparisons (KS, PSI, chi²) of changed columns.
ML debugger Detectors for leakage, train/test contamination, label noise, spurious/id-like features, temporal leakage, class imbalance, outliers, missingness, feature drift and subgroup performance.
Honest model metrics Out-of-fold cross-validation so every number estimates generalisation, not training error.
Pluggable detectors Register your own detector in a few lines (datascope.detectors.register_detector).
Formats CSV, TSV, Parquet, JSON, JSONL + in-memory pandas/polars DataFrame.
Output terminal, json, markdown, html. JSON is machine-parseable; HTML is self-contained and offline.
Behaviour Local-first, deterministic (default seed 42), no telemetry. --seed overrides.

Installation

Requires Python 3.11+.

# from a checkout
pip install -e .

# or with the ML/model-debugging extras (scikit-learn)
pip install -e ".[ml]"

# everything, incl. plotly for optional interactive charts
pip install -e ".[all]"

Core read-only profiling and diffing have no heavy ML dependencies: you get it with a plain pip install ..

CLI sandbox after install:

datascope --help
datascope --version   # datascope 1.0.0

Quickstart

1. Inspect one dataset

datascope inspect ./data/churn_train.csv --format terminal
Dataset: churn_train.csv  (600 rows × 8 columns)
Health:  89 / 100

…per-column profile + issues…

Write a JSON file for scripting:

datascope inspect ./data/churn_train.csv -t churned --format json --output out.json

2. Diff two snapshots

datascope diff ./data/train_v1.csv ./data/train_v2.csv --target churned

Shows how the support/spend/spend_6m distributions changed, marked with the severity of each statistical signal (datascope uses PSI + KS for numerics, JS-divergence/PSI/chi² for categories).

3. Debug a dataset or a model

datascope debug ./data/churn_train.csv --target churned
# PATH is the snapshot you debug; --train/--test (optional) enable
# train/test contamination and temporal-leakage checks against them
datascope debug data/current.csv --train data/train.csv --test data/test.csv --target churned

4. Detect feature drift

datascope drift ./data/train_v1.csv ./data/train_v2.csv

5. Gate a CI pipeline

datascope check ./data/churn_train.csv --target churned --fail-on medium
echo $?   # 0 => clean; 2 => worst finding >= "medium"

Pair with the bundled GitHub Action (.github/workflows/datascope.yml) that calls the same check semantics and fails the workflow when the gate trips.


Python API

Everything the CLI does is a thin wrapper over a small Python library.

from datascope import Dataset, DatasetDiff, MLDebugger, profile_dataset

# 1. Profile & health
resp = profile_dataset("churn.csv", target="churned")
print(resp.score, resp.profile.rows, resp.profile.columns_profiles[0].name)

# 2. Compare two snapshots
d = DatasetDiff("train_v1.csv", "train_v2.csv", target="churned").run()
print(d.overall, len(d.schema), len(d.numeric))

# 3. Model/data debugging
debug = MLDebugger("churn.csv", target="churned").run()
for f in debug.findings:
    print(f.severity.value, f.kind.value, f.title, "->", f.message)

Dataset, DatasetDiff, MLDebugger and profile_dataset accept a path, a CSV/Parquet/JSON/JSONL file, or an in-memory pandas.DataFrame.


Configuration

Add a .datascope.yaml in the working directory (or pass --config) to tune detectors, thresholds, CI gates and the default seed:

# .datascope.yaml
target: churned
seed: 7

diff:
  psi_threshold: 0.3
  ks_pvalue_threshold: 0.01

outliers:
  method: iqr          # iqr | robust_zscore | isolation_forest
  threshold: 4.0

class_balance:
  imbalance_critical: 30.0

model:
  cv_folds: 5
  cv_max_rows: 50000

ablation:
  enabled: true        # opt-in: refits the model repeatedly

ci:
  fail_on: medium
  report_path: datascope-report.md

See docs/configuration.md for every option and its default.


Output formats

Format Usage
terminal Human colour report (default).
json Machine-readable result object (valid JSON on stdout).
markdown Great for PR comments / Jupyter / GitBook.
html Self-contained, offline single-file report (style inlined; no CDN).

Choose with --format/-f or set the output extension (-o out.md → Markdown).

The html report is a single dependency-free document (no CDN, no external fonts) — ideal to attach to CI or email. If you install the optional plotly extra (.[all]), the small datascope.visualization module can also export interactive, self-contained charts for an inspect or drift report from Python; see docs/visualization.md.


Determinism & privacy

  • Default seed is 42 for every command that samples or shuffles. Give --seed to override. Re-running produces identical bytes.
  • All computation is local. Datascope never uploads your data or sends telemetry.
  • JSON/HTML reporters write plain text/files — no external network calls.

Project layout

datascope/
├── src/datascope/
│   ├── cli.py               # typer CLI
│   ├── config.py            # .datascope.yaml parsing/validation
│   ├── dataset.py           # Dataset facade (read + sample + sniff)
│   ├── models.py            # core enums/models (Severity, Finding, …)
│   ├── profiling/           # single-dataset profiling + health scoring
│   ├── diff/                # old-vs-new comparison engine
│   ├── detectors/           # pluggable ML debugger detectors + registry
│   ├── debugging/           # MLDebugger orchestrator / report assembly
│   ├── reporting/           # terminal/json/markdown/html renderers
│   ├── metrics/             # statistical tests (KS/PSI/JS/chi²/AUC/…)
│   ├── loaders/             # CSV/Parquet/JSON/JSONL readers
│   └── utils/               # typing/serialization helpers
├── tests/                   # pytest suite (targets ≥85% branch coverage)
├── docs/                    # configuration, health, architecture, detectors
├── .github/workflows/       # CI that lints + tests + runs datascope check
└── pyproject.toml

Contributing

Contributions are welcome — see CONTRIBUTING (and the security notes in SECURITY). In brief:

pip install -e ".[dev]"
ruff check src && ruff format --check src
mypy src/datascope
pytest --cov=datascope --cov-branch

The detector Protocol (src/datascope/detectors/base.py) plus the registry make it easy to add a new detector without touching the orchestrator.


Documentation


License

Apache-2.0. This is an open-source project; contributions are subject to the terms of the license.

About

A local-first, deterministic CLI + Python toolkit for profiling, diffing, debugging, and monitoring ML datasets.

Topics

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages