Repository navigation
Lift INPI RNE account rejections once INPI is back - #488
Merged
Merged
Conversation
During the INPI outage of 2026-10-09, each server process had its own copy of the rejected accounts: the cache namespace changes on every boot, and two processes booted one second apart held the same rejected accounts under two different namespaces. Each process had to get its own 401 before skipping an account, so INPI received more failed logins than needed, and nothing could lift a rejection for all processes at once. Rejections are now stored in Redis outside the cache namespace, like the ping last success dates, so every process sees the same ones. A deploy no longer lifts them: they last their full 24 hours unless lifted explicitly.
On 2026-10-09, INPI was down and answered 401 on login between 07:15 and 07:55 UTC although our passwords were valid. Both production accounts were rejected for 24 hours, so production kept answering with an INPI RNE maintenance error long after INPI came back, while the ping, which uses its own accounts, was green again. When both accounts are rejected, we now look at the INPI RNE pings: if one succeeded after the last rejection, INPI is back and the accounts are tried again. Otherwise we stay in maintenance as before. Only one request retries per ping success, the others stay in maintenance meanwhile, and rejections are lifted only once a login works: several requests retrying at the same time would multiply the 401 that get our IP banned. Rejection times keep sub-second precision so that a ping earlier in the same second does not count as later. If both production passwords really expire while the ping accounts still work, each new ping success costs one 401 per account. This is rare and shows up in Sentry, which is better than a full day of maintenance after every INPI outage. Covered: no ping success since the rejections, a ping earlier in the same second, recovery, accounts usable again afterwards, a concurrent request during the retry, and expired passwords tried once per ping success. Not covered: Redis errors while retrying, which leave the accounts usable, as before this branch.
skelz0r
force-pushed
the
feature/inpi-401-outage
branch
from
October 9, 2026 09:11
4db8229 to
6485111
Compare
Un3x
approved these changes
Oct 9, 2026
Un3x
left a comment
Contributor
There was a problem hiding this comment.
Je suis médusé par tout ce système en place.
C'est ok, mais j'ai pas envie de faire ca pour chaque FD
Member
Author
Ah bah moi non plus hein 😅 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
On 2026-10-09 INPI was down and answered 401 on login between 07:15 and 07:55 UTC although our passwords were valid. Both production accounts were rejected for 24 hours, so production kept rendering an INPI RNE maintenance error after INPI came back, while the ping, which uses its own accounts, was green again.
The first commit stores the rejections in Redis outside the cache namespace that changes on every boot, so all processes share them: each process no longer needs its own 401 before skipping an account, and a rejection can be lifted for everyone. A deploy no longer lifts them.
The second commit uses the INPI RNE pings when both accounts are rejected: if one succeeded after the last rejection, a single request retries the accounts while the others stay in maintenance, and the rejections are lifted only once a login works. Otherwise we stay in maintenance. If both production passwords really expire while the ping accounts work, each ping success costs one more 401 per account, which is rare and visible in Sentry.