diff --git a/.github/ISSUE_TEMPLATE/bug_report.yml b/.github/ISSUE_TEMPLATE/bug_report.yml index 2d75e54..560a045 100644 --- a/.github/ISSUE_TEMPLATE/bug_report.yml +++ b/.github/ISSUE_TEMPLATE/bug_report.yml @@ -12,12 +12,14 @@ body: Two things answer most reports faster than we can: - - **Zero products?** Check you used a *filtered* category URL, not the - bare `/shopping/kids/items.aspx` hub — the hub carries no product - JSON-LD and correctly returns nothing. See + - **Zero products?** Open the `{out}_page1_debug.html` dump the run + wrote beside itself: a challenge, a sign-in wall or a real page + with no results answers most of these. See [TROUBLESHOOTING.md](../blob/main/TROUBLESHOOTING.md). - - **Wrong currency or localised titles?** Amazon decides both from - your exit IP, not from the URL. See "Geo-redirect" in the README. + - **Wrong currency?** Amazon converts the price to the currency of + your exit IP's country, not the marketplace's — from a European + exit amazon.co.uk and amazon.co.jp both quote EUR. See "The + currency follows the exit IP, not the domain" in the README. - type: dropdown id: engine @@ -66,7 +68,7 @@ body: placeholder: | python3 playwright_scraper.py \ --url "https://www.amazon.com/s?k=wireless+headphones" \ - --pages 1 --out girls + --pages 1 --out headphones validations: required: true @@ -86,7 +88,7 @@ body: attributes: label: What you expected instead placeholder: >- - 96 products, as the README says a category page yields. + 16–30 organic tiles per search page, as the README's measured results say. validations: required: true diff --git a/.github/ISSUE_TEMPLATE/site_changed.yml b/.github/ISSUE_TEMPLATE/site_changed.yml index e4d4d96..efbb4ee 100644 --- a/.github/ISSUE_TEMPLATE/site_changed.yml +++ b/.github/ISSUE_TEMPLATE/site_changed.yml @@ -9,8 +9,12 @@ body: own report because the fix is different from a code bug: something on amazon.com moved. - Useful to know before filing: the parser tries **JSON-LD first**, then a - CSS + URL-pattern fallback. Which one broke narrows the fix a lot. + Useful to know before filing: Amazon publishes no JSON-LD, so the + parser reads the page kind's own data-attribute anchor + (`[data-asin]` on search, `[id^="p13n-asin-index-"]` on best + sellers, `[data-hook="reviewContainer"]` on reviews) and falls back + to the `/dp/{ASIN}` URL pattern, logging a warning when it does. + Which one broke narrows the fix a lot. - type: dropdown id: what_broke @@ -41,9 +45,10 @@ body: description: | Whichever of these you can get. A capture beats a description. - - JSON-LD as the page actually serves it: - `python3 -c "import json,sys,re;h=open('dump.html').read();print([m[:400] for m in re.findall(r'',h,re.S)][:2])"` - - Or open devtools and paste one `Product` node from the `ItemList`. + - The exact bytes the parser was given: rerun with + `--dump-html dump.html` (it writes on success too) and attach it, + with session ids and tokens removed. + - Whether the run log warned that the `/dp/` URL-pattern fallback ran. - For a field problem: the value you got, and the value on the page. render: text validations: diff --git a/CHANGELOG.md b/CHANGELOG.md index 78de6ca..9be3a8e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -126,6 +126,19 @@ one, and when it does the release notes say so first. engine logs the status and does not classify on it). It now reads `http_code`, falling back to `status` only if that is an integer. After the fix, one live call (`--wait-text results` on the README's search URL) answered HTTP 200, upstream 200, 16 rows, status complete. +- **farfetch-scraper leftovers removed from the issue templates, TROUBLESHOOTING, + SECURITY and `diff_runs.py`.** The bug-report tips sent a reader to + Farfetch's `/shopping/kids/items.aspx` hub "which carries no product + JSON-LD", the example wrote `--out girls` and expected "96 products"; the + site-change template said the parser "tries JSON-LD first" and asked for a + JSON-LD dump; TROUBLESHOOTING's 0-rows table named Farfetch's + `-item-.aspx` links and hub; SECURITY named Akamai; `diff_runs.py` + described `source_changed` as "DOM-corrected versus raw JSON-LD" and its + examples used `girls_clothing`. Amazon publishes no JSON-LD, so all of + these now describe the data-attribute anchors, the `/dp/{ASIN}` fallback, + AWS WAF, the `offscreen`/`split`/`detail` price sources and this README's + measured 16–30 tiles per search page. + - **The Scraper API's `x-debug` response header is redacted before it is logged.** `SECURITY.md` names that header as one of three places credentials reach a log unmasked, and the client logged it whole: the API @@ -186,6 +199,9 @@ one, and when it does the release notes say so first. with "Claude Code native binary not found". The channel is pinned rather than a version, so a fixed upstream release needs no edit here. +- `SECURITY.md` said this project has no releases or version tags; it has + both. "Supported versions" now names the latest release and `main`. + ## [0.1.3] — 2026-09-11 ### Fixed diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index c3b9fd9..8e58633 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -117,7 +117,7 @@ exit country, and what you got. Product counts differ by country and by URL, so a bare "worked for me" is not reproducible. -Do not add anything that submits the registration form. This project +Do not add anything that submits a registration or login form. This project deliberately never does, and a captcha token proved valid by creating a real account is not a result worth having. diff --git a/SECURITY.md b/SECURITY.md index 711081a..36631a4 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -55,7 +55,7 @@ Not because these do not matter, but because they belong somewhere else: - **Bypassing Amazon's bot protection.** This scraper drives an ordinary browser and passes challenges the way a browser does. Anything about how - Akamai or reCAPTCHA behave is not a vulnerability in this repository. + AWS WAF or Amazon's own captcha behave is not a vulnerability in this repository. - **The scraper stopped working.** Amazon changing its markup is expected — file it as a normal issue, there is a template for exactly that. - **Anything about 2Captcha's services** — the solver API, the Scraping Browser @@ -79,9 +79,9 @@ Not because these do not matter, but because they belong somewhere else: ## Supported versions -`main` only. This project has no releases or version tags; fixes land on `main` -and you update by pulling. If you are running an old clone, update before -reporting. +The latest release and `main`. Fixes land on `main` first and ship in the +next tagged release (see the Releases page). If you are running an old clone, +update before reporting. ## If you have leaked a key diff --git a/TROUBLESHOOTING.md b/TROUBLESHOOTING.md index 34f55d1..87bda31 100644 --- a/TROUBLESHOOTING.md +++ b/TROUBLESHOOTING.md @@ -11,9 +11,8 @@ answer this in one look. | What the dump shows | Cause | |---|---| | A challenge or "verify you are human" page | Bot management. Use a browser engine (not a plain HTTP fetch), a residential IP, or a remote browser via `--cdp-endpoint`. | -| A real page, prices visible, still 0 rows | The JSON-LD path found nothing and the CSS fallback did not match. Check that product links still match `-item-.aspx`. | +| A real page, prices visible, still 0 rows | The page kind's data-attribute anchor (`[data-asin]`, `[id^="p13n-asin-index-"]`, `[data-hook="reviewContainer"]`) found nothing and the `/dp/{ASIN}` URL-pattern fallback did not match either. See "How it parses" in the README. | | A real page in a different language, prices like `125 €` | Fine — that parses. If rows are still 0, it is not the locale. | -| A near-empty page | The hub URL. `/shopping/kids/items.aspx` has zero products in its JSON-LD; use a filtered category URL. | ## A local Selenium session will not start diff --git a/diff_runs.py b/diff_runs.py index 04d4080..2a69fbb 100644 --- a/diff_runs.py +++ b/diff_runs.py @@ -8,15 +8,15 @@ monitoring and assortment tracking, but that nothing in this repo actually computed. - python3 diff_runs.py --old girls_clothing.2026-09-01.json \\ - --new girls_clothing.2026-09-07.json + python3 diff_runs.py --old headphones.2026-09-01.json \\ + --new headphones.2026-09-07.json Typical use is a scheduled re-run of one of the four scraper engines, kept under a dated filename, diffed against the previous one: - python3 playwright_scraper.py --url "$URL" --out "girls_$(date +%F)" - python3 diff_runs.py --old "girls_$(ls -t girls_*.json | sed -n 2p)" \\ - --new "girls_$(date +%F).json" --out diff.json + python3 playwright_scraper.py --url "$URL" --out "headphones_$(date +%F)" + python3 diff_runs.py --old "headphones_$(ls -t headphones_*.json | sed -n 2p)" \\ + --new "headphones_$(date +%F).json" --out diff.json Four buckets, each keyed on sku: @@ -26,9 +26,9 @@ changed — sku present in both, with a different price, original_price, discount_pct, currency or in_stock source_changed — sku present in both with a different price, but also a - different price_source: one run got the DOM-corrected - figure and the other the raw JSON-LD one, so the two are - not comparable on price. Reported separately because this + different price_source: the two runs read the price from + different DOM nodes (offscreen / split / detail), so the + two are not comparable on price. Reported separately because this says something about our own two snapshots, not about the site — and --fail-on-change deliberately ignores it.