Skip to content

chore: mature run_model_sweep.sh into modular, configurable benchmark suite #38

Description

@AccessiT3ch

Context

scripts/run_model_sweep.sh was authored as a Study 2a research artifact and works, but it is brittle for future studies:

  • Variants are hard-coded as if/elif blocks — adding Study 2b variants requires editing the script
  • No test coverage
  • No --output-dir override — always writes to the study-id-derived path
  • No --timeout override per variant
  • The RAM filter list is a static array, not driven by adaptive_k_selector.py

Acceptance Criteria

  • Variants loaded from a config file (data/sweep-variants.yml or similar) — adding a new variant requires no script edits
  • --timeout and --output-dir CLI flags
  • RAM filter driven by OLLAMA_RAM_LIMIT_GB env var or --ram-limit flag
  • tests/test_run_model_sweep.sh or equivalent (bats, shunit2, or Python subprocess mock)
  • --dry-run output is machine-readable (JSON or structured text)
  • Documented in .github/skills/rag-rapid-research/SKILL.md § 5.2

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions