feat(whisper): timing log, default prompt and hotwords; chart affinity and spread - #16
Merged
Merged
Conversation
…hart scheduling values - whisper: one JSON log line per request (audio duration, decode/VAD vs. inference time, RTF, parameters, prompt/hotwords length only; never content) - whisper: WHISPER_INITIAL_PROMPT / WHISPER_HOTWORDS server-side defaults and a `hotwords` form field; request prompt is appended to the default, request hotwords replace the default - chart: whisper.initialPrompt, whisper.hotwords; optional affinity and topologySpreadConstraints for embedding and whisper (empty by default) - tests: pytest suite for the whisper server (stubbed model), router test for multipart passthrough, CI job for the whisper tests Closes #13
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #13, closes #15
What
whisper_request {...}): audio duration (and after VAD), upload, queue wait, decode+VAD (prep_s), inference (infer_s), total, RTF, model, language, beam size, VAD, file size, and only the length of prompt/hotwords. Prompt, hotwords and transcript text are never logged. No behaviour change.WHISPER_INITIAL_PROMPT/whisper.initialPrompt(empty = off). A requestpromptis appended to it: Whisper weighs the end most, and faster-whisper keeps only the last 223 tokens (get_prompt, 1.2.1), so request text survives and the generic default is cut first. Combined text is capped at 2000 characters (tail kept).hotwordsform field,WHISPER_HOTWORDS/whisper.hotwordsdefault; a request value replaces the default. faster-whisper ignores hotwords only whenprefixis set (never here), so prompt and hotwords combine.affinityandtopologySpreadConstraintsforembeddingandwhisper(empty by default, no change in rendered scheduling). README has a soft anti-affinity example.Rewriteininternal/proxy/proxy.go) only rewrites URL/host headers and streams the multipart body unchanged; new Go test covers prompt/hotwords/language passthrough.Tests
pytest deploy/whisper(stubbed model; parameter forwarding, combination rules, log line without content), new CI jobgo test ./...helm lint,helm templatewith defaults (only the two empty env vars added) and with values setprep_s0.05-0.14 s, the rest is inference.Note: merging changes
deploy/whisper/**, which triggers the Release Models workflow and rebuildswhisper:large-v3-turbo-<YYYY.MM>under the same tag. Deployments withimagePullPolicy: IfNotPresentmust pin the new digest or restart withAlwaysto pick it up.