Skip to content

feat(whisper): timing log, default prompt and hotwords; chart affinity and spread - #16

Merged
eksrha merged 1 commit into
mainfrom
feat/whisper-timing-prompt
Oct 2, 2026
Merged

eksrha merged 1 commit into
mainfrom
feat/whisper-timing-prompt

Conversation

@eksrha

@eksrha eksrha commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Closes #13, closes #15

What

  • Whisper timing log: one JSON line per request (whisper_request {...}): audio duration (and after VAD), upload, queue wait, decode+VAD (prep_s), inference (infer_s), total, RTF, model, language, beam size, VAD, file size, and only the length of prompt/hotwords. Prompt, hotwords and transcript text are never logged. No behaviour change.
  • Default prompt: WHISPER_INITIAL_PROMPT / whisper.initialPrompt (empty = off). A request prompt is appended to it: Whisper weighs the end most, and faster-whisper keeps only the last 223 tokens (get_prompt, 1.2.1), so request text survives and the generic default is cut first. Combined text is capped at 2000 characters (tail kept).
  • Hotwords: new hotwords form field, WHISPER_HOTWORDS / whisper.hotwords default; a request value replaces the default. faster-whisper ignores hotwords only when prefix is set (never here), so prompt and hotwords combine.
  • Chart: optional affinity and topologySpreadConstraints for embedding and whisper (empty by default, no change in rendered scheduling). README has a soft anti-affinity example.
  • Router: no change needed. The reverse proxy (Rewrite in internal/proxy/proxy.go) only rewrites URL/host headers and streams the multipart body unchanged; new Go test covers prompt/hotwords/language passthrough.

Tests

  • pytest deploy/whisper (stubbed model; parameter forwarding, combination rules, log line without content), new CI job
  • go test ./...
  • helm lint, helm template with defaults (only the two empty env vars added) and with values set
  • Local run with the real large-v3-turbo image, 11 s German clip, 4 threads, int8, beam 1, VAD on: RTF 0.76-0.88; decode/resampling prep_s 0.05-0.14 s, the rest is inference.

Note: merging changes deploy/whisper/**, which triggers the Release Models workflow and rebuilds whisper:large-v3-turbo-<YYYY.MM> under the same tag. Deployments with imagePullPolicy: IfNotPresent must pin the new digest or restart with Always to pick it up.

…hart scheduling values

- whisper: one JSON log line per request (audio duration, decode/VAD vs.
  inference time, RTF, parameters, prompt/hotwords length only; never content)
- whisper: WHISPER_INITIAL_PROMPT / WHISPER_HOTWORDS server-side defaults and
  a `hotwords` form field; request prompt is appended to the default,
  request hotwords replace the default
- chart: whisper.initialPrompt, whisper.hotwords; optional affinity and
  topologySpreadConstraints for embedding and whisper (empty by default)
- tests: pytest suite for the whisper server (stubbed model), router test for
  multipart passthrough, CI job for the whisper tests

Closes #13
@eksrha eksrha added the enhancement New feature or request label Oct 2, 2026
@eksrha
eksrha merged commit 9fd62ee into main Oct 2, 2026
3 checks passed
@eksrha
eksrha deleted the feat/whisper-timing-prompt branch October 2, 2026 21:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Whisper: request timing log, default prompt and hotwords Chart: Anti-Affinity bzw. Topology-Spread für Embedding-Replicas

1 participant