Prerequisites
Feature Description
Please implement the feature from https://github.com/GenerelSchwerz/llama.cpp
--live-context-workspace, --no-live-context-workspace
for supported attention caches, grow the compute workspace reservation
with the padded live physical KV extent instead of reserving the full
context up front (default: disabled)
(env: LLAMA_ARG_LIVE_CONTEXT_WORKSPACE)
--phase-aware-workspace, --no-phase-aware-workspace
resize compute workspaces between prompt processing and token
generation; later prompt turns regrow the prompt reservation (default:
disabled)
(env: LLAMA_ARG_PHASE_AWARE_WORKSPACE)
Motivation
It speeds up the token generation at lower context size
Possible Implementation
No response
Prerequisites
Feature Description
Please implement the feature from https://github.com/GenerelSchwerz/llama.cpp
--live-context-workspace, --no-live-context-workspace
for supported attention caches, grow the compute workspace reservation
with the padded live physical KV extent instead of reserving the full
context up front (default: disabled)
(env: LLAMA_ARG_LIVE_CONTEXT_WORKSPACE)
--phase-aware-workspace, --no-phase-aware-workspace
resize compute workspaces between prompt processing and token
generation; later prompt turns regrow the prompt reservation (default:
disabled)
(env: LLAMA_ARG_PHASE_AWARE_WORKSPACE)
Motivation
It speeds up the token generation at lower context size
Possible Implementation
No response