Skip to content

Qwen3.8-27B EXL3 4bpw on Arc Pro B70 (exl3xpu) — 2 results + catalog entries - #49

Merged
jackwsmth merged 3 commits into
labscommunity:mainfrom
SergiioB:exl3-qwen38-results
Sep 28, 2026
Merged

jackwsmth merged 3 commits into
labscommunity:mainfrom
SergiioB:exl3-qwen38-results

Conversation

@SergiioB

Copy link
Copy Markdown
Contributor

Summary

Adds the first EXL3 results for Qwen3.8-27B on Intel Arc Pro B70:

  • results/SergiioB/2026-09-27-…-exl3xpu.json — 150 W cap: decode 17.1 tok/s (client post-first stream rate, median of 9–10 samples, p512/g128), cold 8K prefill ~1,680 tok/s, TTFT ~440 ms, ctx 65,536
  • …-230w.json — stock 230 W cap: sustained 24.6 tok/s (+33–42%), ctx 65,536

Because exl3xpu and exl3-4bpw were not in the catalog, this PR also adds:

  • RUNTIMES: exl3xpu → https://github.com/0xSero/exl3xpu (ExLlamaV3 XPU port; serves via its own vLLM-XPU-derived API surface)
  • QUANTS: exl3-4bpw (EXL3 trellis, ~4 bpw) + the qwen3-8-27b board link
  • supabase/migrations/20260927120000_catalog_exl3_4bpw_exl3xpu.sql mirroring the pattern of earlier catalog migrations

Measurement honesty notes (in the files too)

  • decodeTps is the client-measured post-first-token SSE stream rate. The engine runs native MTP3, so emitted tokens/streamed events exceed 1:1 — the number is a conservative lower bound, not inflated.
  • All samples are fresh measured requests; 230 W TTFT outliers from prefix-cache hits were excluded.
  • Context capability (the point of this route): the model's native 262,144-token context fits on one 32 GB B70 with fp8 KV (pool ~287K); a 261,920-token prompt was answered correctly.
  • Correctness: arithmetic canary clean; notes flag a real codegen caveat (interactive-WebGL battery 2/3).

Validation

  • npm run results:validate — both files pass against the updated catalog
  • npm run catalog:validate — 46 hw / 32 models / 18 quants / 8 runtimes, all resolve

Ingestion may need the catalog migration deployed before the result check passes — happy to split into two PRs if preferred.

… Pro B70

Catalog: new runtime exl3xpu (0xSero/exl3xpu, the ExLlamaV3 XPU port) and
quant exl3-4bpw, linked to the qwen3-8-27b board; seed migration included.

Results: two measured runs from campaign CAM-2026-AUTOROUND-VS-EXL3 —
150 W (median decode 17.1 tok/s post-first stream rate, n=9-10) and
230 W stock cap (24.6). fp8 KV + native MTP3; post-first client rate is a
conservative lower bound under speculative decode. 262,144-token native
context proven on one card; details and caveats in the notes fields.
@vercel

vercel Bot commented Sep 27, 2026

Copy link
Copy Markdown

Someone is attempting to deploy a commit to the Community Labs Team on Vercel.

A member of the Team first needs to authorize it.

@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
- results/SergiioB/2026-09-27-qwen3-8-27b-exl3-4bpw-exl3xpu-230w.json: imported (result 93)
- results/SergiioB/2026-09-27-qwen3-8-27b-exl3-4bpw-exl3xpu.json: imported (result 94)
Merge ingestion completed. Rerunning this workflow will not duplicate these results.

Site / sign up · Workflow details and retry

@SergiioB

Copy link
Copy Markdown
Contributor Author

Trimmed to result files only. The catalog entries (runtime exl3xpu, quant exl3-4bpw) moved to #50 — once that PR is merged and deployed, a maintainer can rerun the ingestion check here.

@jackwsmth
jackwsmth merged commit edc90ad into labscommunity:main Sep 28, 2026
3 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants