Do not fail awaiting cache transactions when the client closes - #595
Open
inakisoriamrf wants to merge 1 commit into
Open
inakisoriamrf wants to merge 1 commit into
inakisoriamrf wants to merge 1 commit into
Conversation
When a cacheable query is in flight, identical queries await its cache transaction. If the client of the first query closed the connection, the query was killed and the transaction was marked as failed. The awaiting queries then got HTTP 500 "[concurrent query failed]" and did not run, although nothing was wrong with the query. Mark the transaction as completed when the query was cancelled because its client closed the connection. The awaiting queries find no cached result and run the query themselves. Timeouts and ClickHouse errors still fail the transaction.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
When a cacheable query is in flight, identical queries do not run: they await the cache transaction of the first query (
AwaitForConcurrentTransaction, up tograce_time).If the client of the first query closes the connection before the response is complete, chproxy kills the query (
KILL QUERY) andcompleteTransactionmarks the transaction as failed. Every query that awaits it then getsHTTP 500with the body[concurrent query failed]and does not run, although the query itself did not fail. A client that leaves early (a closed browser tab, a cancelled request) makes the identical requests of other clients fail.This change marks the transaction as completed when the query was cancelled because its client closed the connection. The awaiting queries find no cached result and run the query themselves, as they already do after a
503or408from ClickHouse. A cancelled response is still not cached. Timeouts (context.DeadlineExceeded) and ClickHouse errors still fail the transaction, so the awaiting queries do not repeat a query that would fail again.Changes:
scope.clientClosedis set inproxyRequestwhen the request ends withcontext.Canceled. In this path the context starts fromcontext.Background(), the configured query timeout usescontext.WithTimeoutand ends withcontext.DeadlineExceeded, and the only other cancellation comes fromlistenToCloseNotify. Socontext.Canceledcan only mean that the client closed the connection.completeTransactioncompletes the transaction whens.clientClosedis true.Pull request type
Please check the type of change your PR introduces:
Checklist
go vet ./...andgolangci-lint runv1.64.8: no issues)go test -race ./...)New test
TestReverseProxy_ClientCloseCompletesCacheTransaction, with a cache that has a grace time:500 "[concurrent query failed] \n". With the change, B runs the query and gets200.500. With the change, it gets200.The test waits for explicit states (the request in flight on the fake server, B not finished and not on the server) instead of fixed delays only. It passed 10 times in a row with
-race, and both cases fail without the change.Does this introduce a breaking change?
Further comments
After a cancelled query, the queries that await it all find a completed transaction and no cached result, so they all run the query. This is the same behaviour as after a
503or408today. Making one of them take over the transaction is a larger change and is out of scope here.An alternative is to leave the transaction failed and let clients retry on
[concurrent query failed]. That moves the problem to every client, and the error message does not tell a cancelled query from a failed one.A related limit that this change does not address: a request that waits in the user queue (
max_queue_time) or inAwaitForConcurrentTransactiondoes not watch its own client. A client that closes during the wait still holds its slot until the wait ends.