Skip to content

Version a - #2

Merged
Atik203 merged 6 commits into
mainfrom
version-A
Aug 5, 2026
Merged

Atik203 merged 6 commits into
mainfrom
version-A

Conversation

@Atik203

@Atik203 Atik203 commented Aug 5, 2026

Copy link
Copy Markdown
Owner

No description provided.

Atik203 added 6 commits August 5, 2026 23:03
- Implement group-aware train/val/test split in `prepare.py` to ensure that both answer variants of a question remain in the same partition. Added integrity report and auto-sampled audit file.
- Update feature extraction process in `extract_features.py` to log the NLI model used for provenance tracking.
- Modify training pipeline in `train_pipeline.py` to include early stopping for XGBoost models and record per-seed metrics for better analysis.
- Introduce a new manifest generator in `make_manifest.py` to document dataset hashes, model versions, and system information for reproducibility.
- Add unit tests in `test_prepare.py` to validate the integrity of group splits and ensure no leakage occurs between train, validation, and test sets.
- Removed obsolete split_indices.npy file.
- Updated make_xgb function to include an early_stopping parameter, allowing for conditional early stopping during model fitting.
- Modified train_seed_models and main functions to utilize the new early_stopping feature, ensuring early stopping is applied only when appropriate.
fix: update import paths for route and parameter types to use dev dir…
Copilot AI lite review requested due to automatic review settings August 5, 2026 18:58
@Atik203
Atik203 merged commit 2dd8767 into main Aug 5, 2026
2 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR tightens Version A reproducibility and integrity by enforcing group-aware dataset splits (to prevent paired-answer leakage), recording provenance/manifest metadata for regenerated artifacts, and updating the web UI/docs to reflect measured results rather than hardcoded claims.

Changes:

  • Implement group-aware item_idx splitting with an integrity report + add pytest coverage to ensure no cross-split leakage.
  • Add provenance/manifest outputs (NLI checkpoint used, per-seed metrics, artifact manifest generation) and update the Colab pipeline accordingly.
  • Update dashboard/about UI and documentation to reflect measured artifacts and corrected versioning/claims.

Reviewed changes

Copilot reviewed 16 out of 18 changed files in this pull request and generated 4 comments.

Show a summary per file
File Description
web/components/risk-gauge.tsx Drive gauge status text/color from API risk label with percent-band fallback.
web/app/dashboard/page.tsx Replace hardcoded “100x cheaper” with cost/latency derived from artifacts.
web/app/about/page.tsx Update feature count + tech stack versions shown on About page.
tests/test_prepare.py Add tests asserting group split is leakage-free and reproducible.
src/models/train_pipeline.py Improve XGBoost device selection, early stopping wiring, and persist per-seed metric rows + NLI model into params.
src/models/make_manifest.py New script to generate an artifacts/results manifest with hashes, versions, and hardware metadata.
src/features/extract_features.py Add batch sizing + device controls; persist NLI checkpoint provenance to JSON.
src/data/prepare.py Switch to group-aware split, save split integrity report, and include group IDs in split indices.
roadmap.md Update roadmap narrative/status to reflect the Version A integrity repair and gating.
README.md Update tables/commands and Colab reference for corrected Version A pipeline.
data/processed/nli_model_used.json Add recorded NLI checkpoint/device/batch provenance output.
data/processed/audit_50_samples.json Add auto-sampled 50-row audit file for manual review.
colab/HaluRISC_Training.ipynb Update Colab pipeline cells to run repaired split protocol + provenance/manifest steps.
blueprint.md Update blueprint status/gates and reflect corrected evidence requirements.
artifacts/split_integrity_report.json Add saved split integrity report artifact (leakage_free = true).
AGENTS.md Update Version A/Version B branch-separation guidance and env keys.
Suppressed comments (1)

web/app/dashboard/page.tsx:139

  • Cost fields are read from latency?.cost_per_1000_usd, but the artifact uses cost_per_1000_predictions_usd (see src/models/eval_efficiency.py). This makes haluriscCost, judgeCost (fallback), and costRatio always null even when the artifact is present.
  // Cost ratio computed from measured artifacts, never hardcoded
  const haluriscCost = latency?.cost_per_1000_usd?.halurisc_local;
  const judgeCost = judge?.cost_per_1000_usd ?? latency?.cost_per_1000_usd?.llm_judge_estimate;
  const costRatio = haluriscCost && judgeCost ? Math.round(judgeCost / haluriscCost) : null;
  const xgbLatencyP50 = latency?.total_per_sample_ms?.p50;

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/data/prepare.py
Comment on lines +78 to +94
def build_integrity_report(df: pd.DataFrame) -> dict:
"""Leakage report: every item_idx (source question) must map to exactly one split."""
per = df.groupby("item_idx")["split"].nunique()
cross = int((per > 1).sum())
report = {
"split": "group_by_item_idx",
"seed": SPLIT_SEED,
"n_groups_total": int(per.size),
"n_groups_per_split": {str(k): int(v) for k, v in df.groupby("split")["item_idx"].nunique().to_dict().items()},
"n_rows_per_split": {str(k): int(v) for k, v in df.groupby("split").size().to_dict().items()},
"label_mean_per_split": {str(k): round(float(v), 4) for k, v in df.groupby("split")["label"].mean().to_dict().items()},
"groups_spanning_multiple_splits": cross,
"leakage_free": cross == 0,
}
if cross > 0:
raise AssertionError(f"Group leakage detected: {cross} item_idx values span multiple splits")
return report
Comment on lines +94 to 98
interface LatencyAnalysis {
n_samples: number;
total_per_sample_ms?: { p50?: number };
cost_per_1000_usd?: { halurisc_local?: number; llm_judge_estimate?: number };
}
Comment thread README.md
Run the **full training pipeline** (feature extraction → XGBoost tuning → calibration → SHAP → RAGTruth validation) on Google Colab with a GPU, then download the artifacts back into this repo:

[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/drive/1aTlrAcIx5FqAaDiYzsLTiMxbMVw1RsRR?usp=sharing)
[![Open In Colab](https://colab.research.google.com/drive/124wjKFVDyZkDNIjs1WgHW8N7vO7XyY6G?usp=sharing)
Comment thread tests/test_prepare.py
Comment on lines +16 to +22
def make_df(n_groups: int = 200, seed: int = 7, missing_halucinated_p: float = 0.1) -> pd.DataFrame:
rng = np.random.default_rng(seed)
rows = []
for i in range(n_groups):
rows.append({"item_idx": i, "label": 0, "x": 1.0})
if rng.random() > missing_halucinated_p:
rows.append({"item_idx": i, "label": 1, "x": 2.0})
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants