Summary
Explore a native memory layer for AgentQuant that helps future search decisions learn from prior experiments while preserving reproducibility and evaluation boundaries. This is a design proposal, not an implementation commitment.
Candidate ideas
1. Separate observations from interpretations
Keep raw run records distinct from diagnoses, generalized findings, and reusable procedures. A model-generated explanation should never be mistaken for measured evidence.
2. Provenance-rich memory records
Attach each memory to its run ID, strategy, asset, regime, data window, harness/config hash, evaluation split, evidence tier, and linked artifacts.
3. Durable source of truth with rebuildable projections
Treat experiment records and artifacts as authoritative. Full-text, embedding, and other retrieval indexes should be disposable and rebuildable, with explicit repair status.
4. Explicit evidence relationships
Represent supports, contradicts, supersedes, reproduces, and derived-from relationships without deleting the historical record. Later evidence should be able to downgrade or supersede an earlier conclusion transparently.
5. Progressive retrieval
Retrieve compact summaries first, filter by structured metadata, and load full traces only for shortlisted candidates. This should reduce context noise and improve auditability.
6. Failure-pattern consolidation
When repeated failures share conditions and causes, produce a reviewable candidate research rule. Preserve the underlying FailureRecords and require traceability back to them.
7. Negative knowledge as a first-class artifact
Store what failed, why it failed, under which conditions, whether the diagnosis was confirmed, and which alternatives were tried afterward.
8. Strict memory scopes
Support filtering by strategy family, asset/market, regime, evaluation protocol, and harness version so superficially similar experiments do not contaminate one another.
9. Temporal visibility guards
During proposal or policy evaluation, expose only memories available before the relevant decision point. Future experiment outcomes must never leak into earlier decisions.
10. Research-state briefs
Generate concise, evidence-linked summaries of what has been tried, recurring failures, promising directions, uncertainty, and the next informative experiment.
Suggested first milestone
Define a typed memory schema and retrieval contract around provenance, evidence tier, scope, temporal visibility, and relationships. Start with existing SQLite-backed research records and FailureRecords; defer sophisticated semantic retrieval until the contract is testable.
Acceptance criteria for a future implementation
- Historical run records remain immutable or append-only.
- Every distilled memory links back to concrete runs and artifacts.
- Retrieval can be filtered by decision timestamp and evaluation split.
- Contradictory findings are both retained and surfaced accurately.
- Failure-derived constraints can be audited and disabled.
- A deterministic fixture demonstrates that future memories are excluded from an earlier decision.
Summary
Explore a native memory layer for AgentQuant that helps future search decisions learn from prior experiments while preserving reproducibility and evaluation boundaries. This is a design proposal, not an implementation commitment.
Candidate ideas
1. Separate observations from interpretations
Keep raw run records distinct from diagnoses, generalized findings, and reusable procedures. A model-generated explanation should never be mistaken for measured evidence.
2. Provenance-rich memory records
Attach each memory to its run ID, strategy, asset, regime, data window, harness/config hash, evaluation split, evidence tier, and linked artifacts.
3. Durable source of truth with rebuildable projections
Treat experiment records and artifacts as authoritative. Full-text, embedding, and other retrieval indexes should be disposable and rebuildable, with explicit repair status.
4. Explicit evidence relationships
Represent supports, contradicts, supersedes, reproduces, and derived-from relationships without deleting the historical record. Later evidence should be able to downgrade or supersede an earlier conclusion transparently.
5. Progressive retrieval
Retrieve compact summaries first, filter by structured metadata, and load full traces only for shortlisted candidates. This should reduce context noise and improve auditability.
6. Failure-pattern consolidation
When repeated failures share conditions and causes, produce a reviewable candidate research rule. Preserve the underlying FailureRecords and require traceability back to them.
7. Negative knowledge as a first-class artifact
Store what failed, why it failed, under which conditions, whether the diagnosis was confirmed, and which alternatives were tried afterward.
8. Strict memory scopes
Support filtering by strategy family, asset/market, regime, evaluation protocol, and harness version so superficially similar experiments do not contaminate one another.
9. Temporal visibility guards
During proposal or policy evaluation, expose only memories available before the relevant decision point. Future experiment outcomes must never leak into earlier decisions.
10. Research-state briefs
Generate concise, evidence-linked summaries of what has been tried, recurring failures, promising directions, uncertainty, and the next informative experiment.
Suggested first milestone
Define a typed memory schema and retrieval contract around provenance, evidence tier, scope, temporal visibility, and relationships. Start with existing SQLite-backed research records and FailureRecords; defer sophisticated semantic retrieval until the contract is testable.
Acceptance criteria for a future implementation