Fetch login-gated Instagram posts, and let corroborated verdicts drain a YouTube queue of dead videos - #5
Merged
Conversation
…n a YouTube queue of dead videos
Two unrelated scrapers had both stopped draining. Instagram cleared 10 of 75
queued posts per run; YouTube cleared none of 190, run after run.
- Instagram scraped public posts only. The yt-dlp legs had run cookie-less
since 2026-07, when attaching cookies made yt-dlp take Instagram's
authenticated web API and that 404'd on every post. Under yt-dlp 2026.8.19
the path works again, and anonymous-only had become the binding constraint:
65 of 75 posts failed every attempt, 50 on Instagram's own ruling ("This
content isn't available to everyone", which matched no rule and churned as a
retryable unknown) and 15 on yt-dlp's "empty media response". Sampled live,
13/13 failed anonymously and 13/13 extracted with the cookies. The scraper
now goes anonymous first and retries once with the session cookies for a
post hidden from logged-out viewers -- immediately, rather than burning
three anonymous attempts on a wall it cannot pass -- and the media leg
follows the metadata leg's auth mode. Anonymous-first keeps the account off
every public post and survives either path breaking again. With the cookies
attached an empty media response means throttling, as it did before.
- A YouTube queue that retries had distilled down to dead videos could never
drain. The metadata leg's ignore_no_formats_error swallowed each refusal, so
a dead video returned an info dict and was saved as an empty placeholder row
(no author, -1 plays, created 2000-01-01); the media leg then answered a
bare "Video unavailable", rightly distrusted since 2026-09-18; and fifteen
of those in a row tripped the permanent-storm guard, whose abort charges no
retry budget. The metadata leg now also asks the tv player client -- the
only one that states why YouTube will not play a video -- and captures the
reason the flag hides. A video the platform has no record of is a failure,
never a placeholder row.
- Corroborated verdicts. The storm guards read a homogeneous run as a broken
session, but a queue of nothing but retries is homogeneous by construction,
so retrying failures guaranteed the guard would trip. A permanent verdict
resting on per-item evidence independent of the error text is now marked
corroborated: it neither extends nor resets a storm run and is pruned even
when the guard trips. YouTube corroborates no-record-plus-permanent-reason,
and a record kept but refused here by region or rights claim (scraped
metadata-only, media leg skipped). A bare "Video unavailable" with the
record intact is never corroborated, so the 2026-09-18 protection stands.
- Classification: "This video is unavailable" was a retryable unknown while
"Video unavailable" was a permanent removal; all nine seen were gone for
good. Content ID blocks get their own permanent "blocked" category, kept
distinct from geo_blocked and removed so another vantage point can single
them out; copyright takedowns read as removals; captcha reads as a bot check.
- Retry strikes: a batch that pruned anything deleted the whole per-platform
sidecar, so one item's success reset a never-succeeding tail's strikes.
Progress now clears only the strikes of the ids it pruned, and ids that
leave the queue drop their media strikes even after an aborted batch.
- Instagram and YouTube are marked residential_ip_only: neither works from a
datacenter IP, and a Cloud Run run would burn the queue and trip guards
whose shared state then holds off the local install. The worker refuses and
the enrichment supervisor skips the queue without charging the plan a stall.
Replayed against the live queue, the next YouTube run takes it from 190 to 2.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two unrelated scrapers had both stopped draining. Instagram cleared 10 of 75
queued posts per run; YouTube cleared none of 190, run after run.
Instagram scraped public posts only. The yt-dlp legs had run cookie-less
since 2026-07, when attaching cookies made yt-dlp take Instagram's
authenticated web API and that 404'd on every post. Under yt-dlp 2026.8.19
the path works again, and anonymous-only had become the binding constraint:
65 of 75 posts failed every attempt, 50 on Instagram's own ruling ("This
content isn't available to everyone", which matched no rule and churned as a
retryable unknown) and 15 on yt-dlp's "empty media response". Sampled live,
13/13 failed anonymously and 13/13 extracted with the cookies. The scraper
now goes anonymous first and retries once with the session cookies for a
post hidden from logged-out viewers -- immediately, rather than burning
three anonymous attempts on a wall it cannot pass -- and the media leg
follows the metadata leg's auth mode. Anonymous-first keeps the account off
every public post and survives either path breaking again. With the cookies
attached an empty media response means throttling, as it did before.
A YouTube queue that retries had distilled down to dead videos could never
drain. The metadata leg's ignore_no_formats_error swallowed each refusal, so
a dead video returned an info dict and was saved as an empty placeholder row
(no author, -1 plays, created 2000-01-01); the media leg then answered a
bare "Video unavailable", rightly distrusted since 2026-09-18; and fifteen
of those in a row tripped the permanent-storm guard, whose abort charges no
retry budget. The metadata leg now also asks the tv player client -- the
only one that states why YouTube will not play a video -- and captures the
reason the flag hides. A video the platform has no record of is a failure,
never a placeholder row.
Corroborated verdicts. The storm guards read a homogeneous run as a broken
session, but a queue of nothing but retries is homogeneous by construction,
so retrying failures guaranteed the guard would trip. A permanent verdict
resting on per-item evidence independent of the error text is now marked
corroborated: it neither extends nor resets a storm run and is pruned even
when the guard trips. YouTube corroborates no-record-plus-permanent-reason,
and a record kept but refused here by region or rights claim (scraped
metadata-only, media leg skipped). A bare "Video unavailable" with the
record intact is never corroborated, so the 2026-09-18 protection stands.
Classification: "This video is unavailable" was a retryable unknown while
"Video unavailable" was a permanent removal; all nine seen were gone for
good. Content ID blocks get their own permanent "blocked" category, kept
distinct from geo_blocked and removed so another vantage point can single
them out; copyright takedowns read as removals; captcha reads as a bot check.
Retry strikes: a batch that pruned anything deleted the whole per-platform
sidecar, so one item's success reset a never-succeeding tail's strikes.
Progress now clears only the strikes of the ids it pruned, and ids that
leave the queue drop their media strikes even after an aborted batch.
Instagram and YouTube are marked residential_ip_only: neither works from a
datacenter IP, and a Cloud Run run would burn the queue and trip guards
whose shared state then holds off the local install. The worker refuses and
the enrichment supervisor skips the queue without charging the plan a stall.
Replayed against the live queue, the next YouTube run takes it from 190 to 2.