Conversation
- Implement group-aware train/val/test split in `prepare.py` to ensure that both answer variants of a question remain in the same partition. Added integrity report and auto-sampled audit file. - Update feature extraction process in `extract_features.py` to log the NLI model used for provenance tracking. - Modify training pipeline in `train_pipeline.py` to include early stopping for XGBoost models and record per-seed metrics for better analysis. - Introduce a new manifest generator in `make_manifest.py` to document dataset hashes, model versions, and system information for reproducibility. - Add unit tests in `test_prepare.py` to validate the integrity of group splits and ensure no leakage occurs between train, validation, and test sets.
…PU utilization and batch processing
- Removed obsolete split_indices.npy file. - Updated make_xgb function to include an early_stopping parameter, allowing for conditional early stopping during model fitting. - Modified train_seed_models and main functions to utilize the new early_stopping feature, ensuring early stopping is applied only when appropriate.
…rtable Colab training, web/dashboard fixes
fix: update import paths for route and parameter types to use dev dir…
There was a problem hiding this comment.
Pull request overview
This PR tightens Version A reproducibility and integrity by enforcing group-aware dataset splits (to prevent paired-answer leakage), recording provenance/manifest metadata for regenerated artifacts, and updating the web UI/docs to reflect measured results rather than hardcoded claims.
Changes:
- Implement group-aware
item_idxsplitting with an integrity report + add pytest coverage to ensure no cross-split leakage. - Add provenance/manifest outputs (NLI checkpoint used, per-seed metrics, artifact manifest generation) and update the Colab pipeline accordingly.
- Update dashboard/about UI and documentation to reflect measured artifacts and corrected versioning/claims.
Reviewed changes
Copilot reviewed 16 out of 18 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
| web/components/risk-gauge.tsx | Drive gauge status text/color from API risk label with percent-band fallback. |
| web/app/dashboard/page.tsx | Replace hardcoded “100x cheaper” with cost/latency derived from artifacts. |
| web/app/about/page.tsx | Update feature count + tech stack versions shown on About page. |
| tests/test_prepare.py | Add tests asserting group split is leakage-free and reproducible. |
| src/models/train_pipeline.py | Improve XGBoost device selection, early stopping wiring, and persist per-seed metric rows + NLI model into params. |
| src/models/make_manifest.py | New script to generate an artifacts/results manifest with hashes, versions, and hardware metadata. |
| src/features/extract_features.py | Add batch sizing + device controls; persist NLI checkpoint provenance to JSON. |
| src/data/prepare.py | Switch to group-aware split, save split integrity report, and include group IDs in split indices. |
| roadmap.md | Update roadmap narrative/status to reflect the Version A integrity repair and gating. |
| README.md | Update tables/commands and Colab reference for corrected Version A pipeline. |
| data/processed/nli_model_used.json | Add recorded NLI checkpoint/device/batch provenance output. |
| data/processed/audit_50_samples.json | Add auto-sampled 50-row audit file for manual review. |
| colab/HaluRISC_Training.ipynb | Update Colab pipeline cells to run repaired split protocol + provenance/manifest steps. |
| blueprint.md | Update blueprint status/gates and reflect corrected evidence requirements. |
| artifacts/split_integrity_report.json | Add saved split integrity report artifact (leakage_free = true). |
| AGENTS.md | Update Version A/Version B branch-separation guidance and env keys. |
Suppressed comments (1)
web/app/dashboard/page.tsx:139
- Cost fields are read from
latency?.cost_per_1000_usd, but the artifact usescost_per_1000_predictions_usd(seesrc/models/eval_efficiency.py). This makeshaluriscCost,judgeCost(fallback), andcostRatioalways null even when the artifact is present.
// Cost ratio computed from measured artifacts, never hardcoded
const haluriscCost = latency?.cost_per_1000_usd?.halurisc_local;
const judgeCost = judge?.cost_per_1000_usd ?? latency?.cost_per_1000_usd?.llm_judge_estimate;
const costRatio = haluriscCost && judgeCost ? Math.round(judgeCost / haluriscCost) : null;
const xgbLatencyP50 = latency?.total_per_sample_ms?.p50;
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Comment on lines
+78
to
+94
| def build_integrity_report(df: pd.DataFrame) -> dict: | ||
| """Leakage report: every item_idx (source question) must map to exactly one split.""" | ||
| per = df.groupby("item_idx")["split"].nunique() | ||
| cross = int((per > 1).sum()) | ||
| report = { | ||
| "split": "group_by_item_idx", | ||
| "seed": SPLIT_SEED, | ||
| "n_groups_total": int(per.size), | ||
| "n_groups_per_split": {str(k): int(v) for k, v in df.groupby("split")["item_idx"].nunique().to_dict().items()}, | ||
| "n_rows_per_split": {str(k): int(v) for k, v in df.groupby("split").size().to_dict().items()}, | ||
| "label_mean_per_split": {str(k): round(float(v), 4) for k, v in df.groupby("split")["label"].mean().to_dict().items()}, | ||
| "groups_spanning_multiple_splits": cross, | ||
| "leakage_free": cross == 0, | ||
| } | ||
| if cross > 0: | ||
| raise AssertionError(f"Group leakage detected: {cross} item_idx values span multiple splits") | ||
| return report |
Comment on lines
+94
to
98
| interface LatencyAnalysis { | ||
| n_samples: number; | ||
| total_per_sample_ms?: { p50?: number }; | ||
| cost_per_1000_usd?: { halurisc_local?: number; llm_judge_estimate?: number }; | ||
| } |
| Run the **full training pipeline** (feature extraction → XGBoost tuning → calibration → SHAP → RAGTruth validation) on Google Colab with a GPU, then download the artifacts back into this repo: | ||
|
|
||
| [](https://colab.research.google.com/drive/1aTlrAcIx5FqAaDiYzsLTiMxbMVw1RsRR?usp=sharing) | ||
| [ |
Comment on lines
+16
to
+22
| def make_df(n_groups: int = 200, seed: int = 7, missing_halucinated_p: float = 0.1) -> pd.DataFrame: | ||
| rng = np.random.default_rng(seed) | ||
| rows = [] | ||
| for i in range(n_groups): | ||
| rows.append({"item_idx": i, "label": 0, "x": 1.0}) | ||
| if rng.random() > missing_halucinated_p: | ||
| rows.append({"item_idx": i, "label": 1, "x": 2.0}) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.