Outcome
Make the output useful enough that researchers share a report and collaborators can inspect or extend the experiment.
Current evidence
Reviewed main 38091c8c on 2026-09-13. Source: src/agent/research_memo.py.
The memo generator already exists and preserves unavailable fields. README says existing-manifest replay remains deferred until episode results are persisted as their own artifact.
Source review only: no new training run, benchmark or external pilot was performed for this issue.
Scope and dependencies
Delivery slice of #28 P3; extend the existing research_memo/episode_report implementation.
Work
- Persist episode outcomes and references to all attempted candidates, configs, memory snapshots, dataset/split hashes and software identities in a versioned bundle.
- Support report regeneration from saved outcomes without rerunning agents; keep this distinct from fresh execution. Include a rerun command and declared data-access requirements.
- Add one user-data ingestion recipe with schema validation, timestamp/order checks and explicit cost assumptions. Keep private/raw data out of shared bundles unless explicitly included.
- Export a self-contained HTML/Markdown decision report: hypothesis, evidence, baseline/candidate diff, unsuccessful candidates, outcome coverage and limitations; share existing renderer semantics.
Acceptance
Outcome
Make the output useful enough that researchers share a report and collaborators can inspect or extend the experiment.
Current evidence
Reviewed main 38091c8c on 2026-09-13. Source: src/agent/research_memo.py.
The memo generator already exists and preserves unavailable fields. README says existing-manifest replay remains deferred until episode results are persisted as their own artifact.
Source review only: no new training run, benchmark or external pilot was performed for this issue.
Scope and dependencies
Delivery slice of #28 P3; extend the existing research_memo/episode_report implementation.
Work
Acceptance