Skip to content

Roadmap: prove budgeted self-improving research with wired harnesses and sealed evaluation #28

Description

@OnePunchMonk

Thesis

Make AgentQuant a reproducible research agent that improves its search policy under a fixed budget, with finance as the first demanding test environment. That is the most interesting part of the project: learning which experiments to run, remembering why candidates failed, and demonstrating transfer to unseen research episodes.

This is a proposed delivery plan, not a claim that the experiments below have succeeded. Reviewed public main at 7c18c308 on 2026-09-12, including implementation, README, benchmark scripts, and open issues. Verification was source inspection plus small isolated probes; I did not run the full suite, a live-market benchmark, or paid LLM evaluations.

What is already worth keeping

  • The zero-key synthetic demo and explicit explanation of its limits.
  • The analyze → propose → backtest → reflect → store loop, structured failure memory, and research workspace.
  • The trailing holdout in run_agent: a meaningful improvement over grading repeated search on the same window.
  • The optional Harnesskit bridge. Keep financial semantics in AgentQuant and make generic experiment/debugging functionality reusable upstream.

Immediate blockers: the experiment must execute what its label says

  1. Epoch configurations are not applied to the agent. _run_epoch receives harness_spec, but calls run_agent with only data, strategy, asset, and trace. Names such as tool-aware, prompt-tuned, and ensemble therefore do not establish that those interventions ran. Add a resolved runtime configuration passed through to every affected component; report requested and effective settings. An unsupported knob must fail clearly.

  2. The epoch “generalization gap” is not a generalization measurement. _compute_metrics computes max(avg_sharpe - best_sharpe, 0) across search results. An isolated execution of this method with search Sharpe 0.6 and holdout Sharpe -0.4 returned gap 0.0. Report search and held-out outcomes separately and define the gap from corresponding windows. Missing holdout results must be unavailable/error, not zero.

  3. The new benchmark does not yet test the headline hypothesis. reproducible_benchmark.py compares one versus three allowed iterations on three synthetic paths and reads sharpe, not holdout_sharpe. More allowed search is a compute intervention, not proof of a better harness. The calls also use persistent memory; isolate/reset state between arms and record an explicit seed for every stochastic component. Enforce offline mode rather than relying on absent environment credentials.

  4. Public claims contradict one another. The README correctly labels GA/DE mock fitness and says forecast accuracy is unreported, yet still contains “Live Results,” “improvements are real,” “86% validated,” and incompatible test counts. Replace the repeated metrics with one generated evidence table: fixture/demo, measured historical experiment, or unverified legacy result. Increasing tool calls from zero to eight is not an 8× efficiency improvement.

  5. Trace export currently invents execution semantics. The bridge maps every hypothesize event to llm_call, including grid-only proposals; an isolated probe confirms this. Other stages become tool calls, usage defaults to zero, and completion is unconditional. Represent internal stages separately; unknown usage must remain unknown, and final task outcomes must come from the actual run.

These refine existing #19–#21 rather than asking for another parallel verification system. In particular, simply rerunning the six-epoch script does not resolve #19 until configurations and metrics are correct. For #20, deterministic fixtures belong in ordinary CI; stochastic live runs should produce versioned evidence, not be required to reproduce a fixed Sharpe within 1% on every PR.

Build in this order

P0 — Make one experiment trustworthy

  • Wire an immutable, versioned harness configuration through proposal generation, tool admission, memory, prompts, and stopping policy.
  • With scripted clients, demonstrate that disabling tools results in zero external calls, changing a prompt changes the submitted prompt, and each supported mutation changes the intended behavior.
  • Persist run ID, parent policy, data hash, time boundary, configuration hash, memory snapshot ID, seeds, attempted candidates, actual fallback path, failures, and costs.
  • Separate passed_quality_gate, budget_exhausted, no_valid_candidate, and execution_failed. The current reflect path labels the best available result accepted after the iteration limit even when it misses the threshold.
  • Repair and reconcile README/results claims and link each measured number to its manifest and command.

Done: two configurations produce verifiably different intended behavior; their report remains honest when the backend fails or the holdout is missing.

P1 — Establish a fair search benchmark

Use chronological outer episodes. Inside each episode, proposal search and policy selection see development data only; final grading uses a sealed later window. Reset memory between comparative arms, or initialize both from the same frozen snapshot. For continual-learning tests, allow only information available before each episode.

Historical web research needs an as-of corpus. Today's search results and model pretraining can contain the future of a historical backtest; a price lookback guard does not solve that. Label residual contamination risk and use prospective paper research for stronger evidence.

Compare:

Arm What it answers
Fixed strategy / buy-and-hold where meaningful Is the market task informative?
Random and grid search Does the agent add value over cheap search?
Frozen agent Does adaptation add anything?
Frozen agent plus memory Is retrieval the source of the gain?
Bounded policy mutation Does changing the research process transfer?

Hold candidate/backtest budgets fixed and report full LLM/tool/time cost; additionally compare quality at matched dollar budgets where feasible. Keep failed attempts in the denominator. Report held-out net returns, drawdown, turnover, search efficiency, and uncertainty; resample appropriate time blocks or episode clusters rather than treating correlated bars as independent.

  • Frozen data/splits and explicit transaction-cost assumptions.
  • At least three search seeds per arm as an initial variance check; expand based on uncertainty, not a predetermined success claim.
  • All attempted hypotheses retained for selection-bias analysis.
  • Leakage sentinels, future-dated memory, and missing outcomes tested.
  • Publish a result even if the adaptive agent ties or loses.

P2 — Implement actual bounded self-improvement

Keep two loops distinct: inner loop searches strategies; outer loop changes how the researcher searches.

Start with one mutation family—proposal prompt or allocation between existing generators. A candidate includes parent ID, patch, diagnosis, expected benefit, evaluation budget, and rollback reference. Evaluate on development episodes, select on validation, then evaluate the frozen selected policy once on final episodes. Protect against repeated reuse of the final holdout.

A GEPA-style reflective proposal mechanism is a useful experimental baseline, not a reason to immediately build a general autonomous code editor. Compare it with random mutations under the same budget. Add memory-disabled and shuffled-memory ablations to determine whether useful evidence, rather than extra context, drives gains.

Promotion criterion: a predeclared practically meaningful improvement at fixed resources, without unacceptable drawdown/regression on protected slices. Inconclusive evidence keeps the incumbent. Persisting a new config alone is not self-improvement.

P3 — Turn it into a compelling research workspace

Show one complete research episode: hypothesis → evidence available at the time → experiment → rejection/acceptance → policy change → fresh result. Let users inspect unsuccessful candidates, compare policies, and export a reproducible research memo. Add prospective paper research only after the experiment contract is stable.

What to remove or defer

Research that informs these choices

  • GEPA (2025): uses trajectory feedback to propose and compare prompt updates. Borrow the bounded reflective mutation experiment; its published gains do not establish gains on financial research.
  • R&D-Agent: a relevant comparator for automated factor/model research. Differentiate on auditable search-policy improvement rather than the existence of multiple agents.
  • TradingAgents (2024): establishes multi-role financial agents as an existing approach; another analyst/critic role is not by itself a new contribution.
  • Bailey and López de Prado, Deflated Sharpe Ratio (2014): motivates logging the complete search history and accounting for selection effects. Its assumptions and effective trial count must be stated; it does not repair leakage.
  • Anthropic, Demystifying evals for AI agents (2026): distinguishes task outcomes, trials, graders, and traces. Apply this to separate successful research from merely completing the loop.

First release worth showing

One command compares frozen versus adaptive research policies, reproduces every reported result, shows a real failure and its attempted correction, and reports whether the correction transfers to untouched episodes. A credible negative result here is more valuable than another unsupported improvement chart.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions