Skip to content

Fetch login-gated Instagram posts, and let corroborated verdicts drain a YouTube queue of dead videos - #5

Merged
pwikstrom merged 1 commit into
mainfrom
claude/scraper-failures-instagram-youtube-f90476
Sep 21, 2026
Merged

pwikstrom merged 1 commit into
mainfrom
claude/scraper-failures-instagram-youtube-f90476

Conversation

@pwikstrom

Copy link
Copy Markdown
Owner

Two unrelated scrapers had both stopped draining. Instagram cleared 10 of 75
queued posts per run; YouTube cleared none of 190, run after run.

  • Instagram scraped public posts only. The yt-dlp legs had run cookie-less
    since 2026-07, when attaching cookies made yt-dlp take Instagram's
    authenticated web API and that 404'd on every post. Under yt-dlp 2026.8.19
    the path works again, and anonymous-only had become the binding constraint:
    65 of 75 posts failed every attempt, 50 on Instagram's own ruling ("This
    content isn't available to everyone", which matched no rule and churned as a
    retryable unknown) and 15 on yt-dlp's "empty media response". Sampled live,
    13/13 failed anonymously and 13/13 extracted with the cookies. The scraper
    now goes anonymous first and retries once with the session cookies for a
    post hidden from logged-out viewers -- immediately, rather than burning
    three anonymous attempts on a wall it cannot pass -- and the media leg
    follows the metadata leg's auth mode. Anonymous-first keeps the account off
    every public post and survives either path breaking again. With the cookies
    attached an empty media response means throttling, as it did before.

  • A YouTube queue that retries had distilled down to dead videos could never
    drain. The metadata leg's ignore_no_formats_error swallowed each refusal, so
    a dead video returned an info dict and was saved as an empty placeholder row
    (no author, -1 plays, created 2000-01-01); the media leg then answered a
    bare "Video unavailable", rightly distrusted since 2026-09-18; and fifteen
    of those in a row tripped the permanent-storm guard, whose abort charges no
    retry budget. The metadata leg now also asks the tv player client -- the
    only one that states why YouTube will not play a video -- and captures the
    reason the flag hides. A video the platform has no record of is a failure,
    never a placeholder row.

  • Corroborated verdicts. The storm guards read a homogeneous run as a broken
    session, but a queue of nothing but retries is homogeneous by construction,
    so retrying failures guaranteed the guard would trip. A permanent verdict
    resting on per-item evidence independent of the error text is now marked
    corroborated: it neither extends nor resets a storm run and is pruned even
    when the guard trips. YouTube corroborates no-record-plus-permanent-reason,
    and a record kept but refused here by region or rights claim (scraped
    metadata-only, media leg skipped). A bare "Video unavailable" with the
    record intact is never corroborated, so the 2026-09-18 protection stands.

  • Classification: "This video is unavailable" was a retryable unknown while
    "Video unavailable" was a permanent removal; all nine seen were gone for
    good. Content ID blocks get their own permanent "blocked" category, kept
    distinct from geo_blocked and removed so another vantage point can single
    them out; copyright takedowns read as removals; captcha reads as a bot check.

  • Retry strikes: a batch that pruned anything deleted the whole per-platform
    sidecar, so one item's success reset a never-succeeding tail's strikes.
    Progress now clears only the strikes of the ids it pruned, and ids that
    leave the queue drop their media strikes even after an aborted batch.

  • Instagram and YouTube are marked residential_ip_only: neither works from a
    datacenter IP, and a Cloud Run run would burn the queue and trip guards
    whose shared state then holds off the local install. The worker refuses and
    the enrichment supervisor skips the queue without charging the plan a stall.

Replayed against the live queue, the next YouTube run takes it from 190 to 2.

…n a YouTube queue of dead videos

Two unrelated scrapers had both stopped draining. Instagram cleared 10 of 75
queued posts per run; YouTube cleared none of 190, run after run.

- Instagram scraped public posts only. The yt-dlp legs had run cookie-less
  since 2026-07, when attaching cookies made yt-dlp take Instagram's
  authenticated web API and that 404'd on every post. Under yt-dlp 2026.8.19
  the path works again, and anonymous-only had become the binding constraint:
  65 of 75 posts failed every attempt, 50 on Instagram's own ruling ("This
  content isn't available to everyone", which matched no rule and churned as a
  retryable unknown) and 15 on yt-dlp's "empty media response". Sampled live,
  13/13 failed anonymously and 13/13 extracted with the cookies. The scraper
  now goes anonymous first and retries once with the session cookies for a
  post hidden from logged-out viewers -- immediately, rather than burning
  three anonymous attempts on a wall it cannot pass -- and the media leg
  follows the metadata leg's auth mode. Anonymous-first keeps the account off
  every public post and survives either path breaking again. With the cookies
  attached an empty media response means throttling, as it did before.

- A YouTube queue that retries had distilled down to dead videos could never
  drain. The metadata leg's ignore_no_formats_error swallowed each refusal, so
  a dead video returned an info dict and was saved as an empty placeholder row
  (no author, -1 plays, created 2000-01-01); the media leg then answered a
  bare "Video unavailable", rightly distrusted since 2026-09-18; and fifteen
  of those in a row tripped the permanent-storm guard, whose abort charges no
  retry budget. The metadata leg now also asks the tv player client -- the
  only one that states why YouTube will not play a video -- and captures the
  reason the flag hides. A video the platform has no record of is a failure,
  never a placeholder row.

- Corroborated verdicts. The storm guards read a homogeneous run as a broken
  session, but a queue of nothing but retries is homogeneous by construction,
  so retrying failures guaranteed the guard would trip. A permanent verdict
  resting on per-item evidence independent of the error text is now marked
  corroborated: it neither extends nor resets a storm run and is pruned even
  when the guard trips. YouTube corroborates no-record-plus-permanent-reason,
  and a record kept but refused here by region or rights claim (scraped
  metadata-only, media leg skipped). A bare "Video unavailable" with the
  record intact is never corroborated, so the 2026-09-18 protection stands.

- Classification: "This video is unavailable" was a retryable unknown while
  "Video unavailable" was a permanent removal; all nine seen were gone for
  good. Content ID blocks get their own permanent "blocked" category, kept
  distinct from geo_blocked and removed so another vantage point can single
  them out; copyright takedowns read as removals; captcha reads as a bot check.

- Retry strikes: a batch that pruned anything deleted the whole per-platform
  sidecar, so one item's success reset a never-succeeding tail's strikes.
  Progress now clears only the strikes of the ids it pruned, and ids that
  leave the queue drop their media strikes even after an aborted batch.

- Instagram and YouTube are marked residential_ip_only: neither works from a
  datacenter IP, and a Cloud Run run would burn the queue and trip guards
  whose shared state then holds off the local install. The worker refuses and
  the enrichment supervisor skips the queue without charging the plan a stall.

Replayed against the live queue, the next YouTube run takes it from 190 to 2.
@pwikstrom
pwikstrom merged commit 67a17b8 into main Sep 21, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant