Local LLM benchmark suite comparing models served via LM Studio OpenAI-compatible API. Designed for Apple Silicon, but works with any LM Studio endpoint.
# 1. Clone & install
git clone https://github.com/omar16100/llm-benchmark
cd llm-benchmark
uv sync
# 2. Load your models in LM Studio
lms server start -p 1234
lms load your-model-1 --identifier model-1
lms load your-model-2 --identifier model-2
# 3. Edit MODELS dict in run_bench.py to match your loaded identifiers
# 4. Run benchmark
uv run python run_bench.py
# 5. (Optional) Claude-as-judge scoring: shells out to the `claude` CLI
# (`claude -p`), which uses that CLI's own login
uv run python judge_claude.pyRuns 26 prompts across 6 categories against each model configured in MODELS (run_bench.py). Captures:
- Quality: exact match, unit tests, constraint checks, tool-call validation, Claude-as-judge
- Performance: time-to-first-token, generation time, total time, approximate tokens/second
- Pairwise judging:
judge_claude.pycompares two models blind, per prompt (A, B, or tie)
| Code | Category | Count | Scoring |
|---|---|---|---|
| R | reasoning | 5 | exact, keyword |
| C | coding | 5 | unit tests, judge |
| M | math | 4 | exact |
| I | instruction | 4 | exact, constraint |
| W | creative | 4 | judge, constraint |
| T | tool_use | 4 | tool trace |
The table below is generated from data/runs_20apr2026.csv, a copy of a local results/runs.csv (705 rows) exported with scripts/runs_data.py (the only change is that absolute model paths are reduced to their directory names):
uv run python scripts/runs_data.py summary data/runs_20apr2026.csv| model_label | quant | valid/total | scored | mean /5 | reasoning | coding | math | instruction | creative | tool_use | median total_s |
|---|---|---|---|---|---|---|---|---|---|---|---|
| supergemma4_26b_mlx4_v2 | MLX-4bit | 54/78 | 48 | 4.41 | 5.00 | 5.00 | 3.33 | 5.00 | 4.78 | 3.75 | 18.0 |
| glm_51_mlx_36bit | MLX-3.6bit | 75/78 | 69 | 4.37 | 4.00 | 5.00 | 3.25 | 5.00 | 5.00 | 4.17 | 131.3 |
| minimax_m27_mlx4_mxfp4 | MLX-4bit-mxfp4 | 69/78 | 63 | 4.16 | 2.80 | 5.00 | 3.33 | 5.00 | 5.00 | 4.06 | 18.9 |
| qwen35_122b_a10b_q8 | GGUF-Q8_0 | 72/78 | 66 | 4.09 | 3.00 | 5.00 | 2.50 | 5.00 | 5.00 | 3.76 | 61.4 |
| minimax_m27_mlx4_mxfp4_mlxlm | MLX-4bit-mxfp4 | 60/78 | 54 | 3.73 | 0.60 | 5.00 | 3.33 | 5.00 | 4.89 | 2.81 | 15.6 |
| qwen35_122b_a10b_q4km | GGUF-Q4_K_M | 51/78 | 45 | 3.58 | 2.50 | 5.00 | 0.00 | 5.00 | 4.56 | 3.76 | 41.1 |
| qwen35_397b_mlx4 | MLX-Q4 | 57/78 | 51 | 2.78 | 2.50 | 5.00 | 1.00 | 5.00 | 4.78 | 0.00 | 41.7 |
| glm_47_flash_q4km_32k | GGUF-Q4_K_M | 30/78 | 27 | 2.39 | 2.75 | n/a | n/a | 0.00 | 3.67 | 2.50 | 48.1 |
| glm_47_flash_q4km | GGUF-Q4_K_M | 24/78 | 21 | 2.14 | n/a | n/a | n/a | 0.00 | 3.33 | 2.50 | 8.9 |
| qwen35_27b_claude_opus_distilled_q8 | GGUF-Q8_0 | 3/3 | 3 | 0.00 | n/a | n/a | 0.00 | n/a | n/a | n/a | 50.3 |
How to read it:
mean /5and the category columns averagescore_rawover valid rows that have a programmatic score. Rows scoredneeds_judgehave none, soscoredcan be lower thanvalid;n/ameans no scored rows in that category.- A row is invalid when the response errored, was empty, or was length-truncated on a non-creative prompt (
is_invalid_resultinrun_bench.py). A full run is 26 prompts x 3 repeats = 78 rows;qwen35_27b_claude_opus_distilled_q8is a 3-row smoke run. median total_sis the median of the per-rowtotal_sover valid rows (wall time of one scored run; for tool-use cases, the sum over its turns). Throughput is not shown becausetok_per_sis not comparable across these runs: rows written before the runner started counting reasoning characters (commit fd9c240) under-count thinking models (see docs/results.md).- Models ran on different runtimes (LM Studio,
mlx_lm.server) and quantizations, so this is a deployment comparison, not an architecture ranking (see Fair Comparison Caveats below). Hardware is not recorded in the run data.
Per-model notes and run history: docs/results.md.
The earlier Gemma 4 31B vs Qwen 3.5 27B run (5 Apr 2026) has no results file; what survives in the LM Studio server logs and the harness log is recovered, with every unrecoverable number listed, in docs/27092026_recovered_gemma4_qwen35_run.md.
cases.json 26 prompts with metadata, expected answers, constraints
run_bench.py Benchmark runner: streams to LM Studio, scores, writes CSV
judge_claude.py Blind pairwise Claude-as-judge scorer
bench_longctx.py Long-context needle + prefill-throughput eval (any OpenAI-compatible server)
bench_long_prompt.py Prompt-length sweep; prefill/decode tok/s from server timings
bench_common.py Shared OpenAI-compatible chat call + server timing extraction
bench_niah_mlx.py 8-needle NIAH ladder, in-process mlx-lm (templated haystack)
*_niah_client.py Same NIAH scoring for raw prompts on mlx-lm, llama-server,
oMLX, mlx-dspark, and a glm5_next mlx-vlm fork
niah_haystack.py Token-accurate 8-needle haystack builder
make_1m_prompt.py Build a ~1M-token NIAH prompt from a source prompt's needles
(evenly spaced depths)
lcb_local_runner.py LiveCodeBench against a local server, with lcb_guards.py
macmon_clamp_report.py GPU clock / RAM summary for a macmon trace
scripts/runs_data.py Export a publishable runs CSV and generate the sample results table
data/ Published run data (runs_20apr2026.csv)
tests/ Unit tests: 142 run without a model, server, or GPU; the DeepEval
tests need a live endpoint, the NIAH haystack tests need tiktoken
plus a local tokenizer, and the mlx tests need mlx-lm, so they
skip otherwise
results/ Local outputs, gitignored
runs.csv Per-run metrics (timing, scores)
transcripts.jsonl Full responses + reasoning traces
The runner uses raw httpx streaming to capture reasoning (reasoning_content from LM Studio, reasoning from mlx_lm.server) separately from content. Every scored request (including each tool-use turn) uses MAX_RESPONSE_TOKENS = 32768 because reasoning tokens share the budget with the final answer; warmup calls use min(case max_tokens, 64). The max_tokens column in runs.csv records the case's nominal budget from cases.json.
bench_longctx.py measures long-context retrieval and prefill throughput against any OpenAI-compatible endpoint (llama-server, LM Studio, vLLM). It builds a haystack of roughly N tokens, inserts a unique needle at each requested depth, asks for it back, and checks exact recall. Prefill and decode tok/s come from the server timings block when present (true server-side throughput), so client wall clock is never reported as prefill.
uv run python bench_longctx.py \
--base-url http://127.0.0.1:8081 --model glm-5.2 \
--target-tokens 2000 8000 32000 --depths 25 50 90 \
--no-thinking --stream --json out.json --csv out.csvEach result row records: target_tokens, depth_pct, prompt_tokens, prefill_tps, decode_tps, ttft_s (with --stream), end_to_end_s, recall (PASS/FAIL), and a free-text server label. Use --no-thinking for GLM and Qwen reasoning models (it sends chat_template_kwargs.enable_thinking=false); omit it for servers that reject unknown template kwargs. bench_common.py holds the shared endpoint call and timing extraction, reused by bench_long_prompt.py.
By default each cell disables server prompt caching (cache_prompt: false, llama.cpp) so prefill is measured cold and is comparable across cells; pass --cache-prompt to keep caching on. prompt_tokens is the full context size from usage, not the server-evaluated subset. The base URL is accepted with or without a /v1 suffix and falls back to BENCH_BASE_URL / BENCH_API_KEY (the same contract as run_bench.py). Full reference: docs/longctx_bench.md.
A second family of long-context tools runs an 8-needle needle-in-a-haystack prompt at 256K tokens and beyond through several local engines. They share the needle format and record retrieval and association scores, a decode figure (a least-squares slope over per-token times for most clients), and guard fields: truncated when the engine saw a different number of prompt tokens than were built (bench_niah_mlx.py refuses to score such a run), and answer_inconclusive when a reasoning model hit the generation cap inside an unterminated <think> block. Which client has which guard, and where the scorers and decode figures differ, is tabulated in the docs. Engines covered: mlx-lm in-process (bench_niah_mlx.py with a templated haystack, mlx_raw_niah_client.py with a raw prompt file), llama-server, oMLX, mlx-dspark, and a patched mlx-vlm fork for glm5_next models. All default endpoints are on 127.0.0.1; local paths come from flags or environment variables (NIAH_MODEL_DIR, GLM53_FLASH_MLX_DIR, NIAH_256K_PROMPT, LCB_DIR, NEMOTRON_MLX8_DIR, BONSAI_FAST_DIR).
# same raw prompt file through two engines
uv run --with mlx-lm python mlx_raw_niah_client.py --model-dir /path/to/model \
--prompt-file niah_256k_prompt.txt --out mlx_raw.json
uv run python llamacpp_niah_client.py --prompt-file niah_256k_prompt.txt --out llamacpp.jsonThese harnesses were built for a 256K cross-engine comparison kept elsewhere; no results from it are published in this repo. Details, the prompt format, how to build a prompt file, and the main flags: docs/niah_harnesses.md.
Edit cases.json:
{
"id": "R6",
"category": "reasoning",
"scoring": "exact",
"temperature": 0.0,
"max_tokens": 256,
"prompt": "Your question here",
"expected": "expected answer"
}Scoring types:
exact: requiresexpectedfield, checks exact match (5pt) or substring (3pt)judge_keyword: requiresexpected_keywordslistunit_tests: requirestest_code(Python assertions)constraint_check: requiresconstraintsdict (word_count, required_words, forbidden_chars, line_count, word_range, etc.)tool_trace_exact/tool_trace_judge: for tool_use cases with mock responsesjudge: deferred to the Claude judge (score_typeneeds_judge)judge_constraint: constraint score whenconstraintsis present, otherwise deferred to the Claude judge
Edit MODELS dict in run_bench.py:
MODELS = {
"my_model": {
"served_model": "lmstudio-identifier",
"quant": "Q4_K_M",
"thinking": False,
},
"thinking_model": {
"served_model": "another-id",
"quant": "Q8_0",
"thinking": True,
},
}uv run pytest tests/ -v- 2 warmup calls discarded, 3 scored repeats per prompt per model
- Deterministic tasks:
temperature=0, top_p=1, seed=42 - Creative tasks:
temperature=0.8, top_p=0.95, seeds 41/42/43 - Concurrency: 1 request at a time
This is a deployment comparison, not pure architecture. Comparing models at different quantizations (bf16, Q8_0, Q4, MLX 4-bit) and runtimes folds together:
- Model quality
- Quantization precision
- Memory bandwidth
- Runtime stack
Reported conclusions should be framed as "practical winner on this hardware," not "model A is better than model B."
MIT