Skip to content
#

bigcodebench

Here are 4 public repositories matching this topic...

Language: All
Filter by language

Benchmark suite for evaluating LLMs and SLMs on coding and SE tasks. Features HumanEval, MBPP, SWE-bench, and BigCodeBench with an interactive Streamlit UI. Supports cloud APIs (OpenAI, Anthropic, Google) and local models via Ollama. Tracks pass rates, latency, token usage, and costs.

  • Updated Apr 23, 2026
  • Python

Resample or reroute after a weak-verifier stop? Pre-registered measurements of recoverable stopping debt on MBPP+, a two-sided action-support gate on BigCodeBench that fails closed, a LiveCodeBench observability ladder, and the exchangeable-actions reference showing realized-maximum gaps carry no selector signal. Artifacts for arXiv:2607.08665v3.

  • Updated Sep 4, 2026
  • Python

Add this topic to your repo

To associate your repository with the bigcodebench topic, visit your repo's landing page and select "manage topics."

Learn more