diff --git a/.github/ISSUE_TEMPLATE/bug_report.yml b/.github/ISSUE_TEMPLATE/bug_report.yml
index 2d75e54..560a045 100644
--- a/.github/ISSUE_TEMPLATE/bug_report.yml
+++ b/.github/ISSUE_TEMPLATE/bug_report.yml
@@ -12,12 +12,14 @@ body:
Two things answer most reports faster than we can:
- - **Zero products?** Check you used a *filtered* category URL, not the
- bare `/shopping/kids/items.aspx` hub — the hub carries no product
- JSON-LD and correctly returns nothing. See
+ - **Zero products?** Open the `{out}_page1_debug.html` dump the run
+ wrote beside itself: a challenge, a sign-in wall or a real page
+ with no results answers most of these. See
[TROUBLESHOOTING.md](../blob/main/TROUBLESHOOTING.md).
- - **Wrong currency or localised titles?** Amazon decides both from
- your exit IP, not from the URL. See "Geo-redirect" in the README.
+ - **Wrong currency?** Amazon converts the price to the currency of
+ your exit IP's country, not the marketplace's — from a European
+ exit amazon.co.uk and amazon.co.jp both quote EUR. See "The
+ currency follows the exit IP, not the domain" in the README.
- type: dropdown
id: engine
@@ -66,7 +68,7 @@ body:
placeholder: |
python3 playwright_scraper.py \
--url "https://www.amazon.com/s?k=wireless+headphones" \
- --pages 1 --out girls
+ --pages 1 --out headphones
validations:
required: true
@@ -86,7 +88,7 @@ body:
attributes:
label: What you expected instead
placeholder: >-
- 96 products, as the README says a category page yields.
+ 16–30 organic tiles per search page, as the README's measured results say.
validations:
required: true
diff --git a/.github/ISSUE_TEMPLATE/site_changed.yml b/.github/ISSUE_TEMPLATE/site_changed.yml
index e4d4d96..efbb4ee 100644
--- a/.github/ISSUE_TEMPLATE/site_changed.yml
+++ b/.github/ISSUE_TEMPLATE/site_changed.yml
@@ -9,8 +9,12 @@ body:
own report because the fix is different from a code bug: something on
amazon.com moved.
- Useful to know before filing: the parser tries **JSON-LD first**, then a
- CSS + URL-pattern fallback. Which one broke narrows the fix a lot.
+ Useful to know before filing: Amazon publishes no JSON-LD, so the
+ parser reads the page kind's own data-attribute anchor
+ (`[data-asin]` on search, `[id^="p13n-asin-index-"]` on best
+ sellers, `[data-hook="reviewContainer"]` on reviews) and falls back
+ to the `/dp/{ASIN}` URL pattern, logging a warning when it does.
+ Which one broke narrows the fix a lot.
- type: dropdown
id: what_broke
@@ -41,9 +45,10 @@ body:
description: |
Whichever of these you can get. A capture beats a description.
- - JSON-LD as the page actually serves it:
- `python3 -c "import json,sys,re;h=open('dump.html').read();print([m[:400] for m in re.findall(r'',h,re.S)][:2])"`
- - Or open devtools and paste one `Product` node from the `ItemList`.
+ - The exact bytes the parser was given: rerun with
+ `--dump-html dump.html` (it writes on success too) and attach it,
+ with session ids and tokens removed.
+ - Whether the run log warned that the `/dp/` URL-pattern fallback ran.
- For a field problem: the value you got, and the value on the page.
render: text
validations:
diff --git a/CHANGELOG.md b/CHANGELOG.md
index 78de6ca..9be3a8e 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -126,6 +126,19 @@ one, and when it does the release notes say so first.
engine logs the status and does not classify on it). It now reads
`http_code`, falling back to `status` only if that is an integer. After the fix, one live call (`--wait-text results` on the README's search URL) answered HTTP 200, upstream 200, 16 rows, status complete.
+- **farfetch-scraper leftovers removed from the issue templates, TROUBLESHOOTING,
+ SECURITY and `diff_runs.py`.** The bug-report tips sent a reader to
+ Farfetch's `/shopping/kids/items.aspx` hub "which carries no product
+ JSON-LD", the example wrote `--out girls` and expected "96 products"; the
+ site-change template said the parser "tries JSON-LD first" and asked for a
+ JSON-LD dump; TROUBLESHOOTING's 0-rows table named Farfetch's
+ `-item-.aspx` links and hub; SECURITY named Akamai; `diff_runs.py`
+ described `source_changed` as "DOM-corrected versus raw JSON-LD" and its
+ examples used `girls_clothing`. Amazon publishes no JSON-LD, so all of
+ these now describe the data-attribute anchors, the `/dp/{ASIN}` fallback,
+ AWS WAF, the `offscreen`/`split`/`detail` price sources and this README's
+ measured 16–30 tiles per search page.
+
- **The Scraper API's `x-debug` response header is redacted before it is
logged.** `SECURITY.md` names that header as one of three places
credentials reach a log unmasked, and the client logged it whole: the API
@@ -186,6 +199,9 @@ one, and when it does the release notes say so first.
with "Claude Code native binary not found". The channel is pinned rather
than a version, so a fixed upstream release needs no edit here.
+- `SECURITY.md` said this project has no releases or version tags; it has
+ both. "Supported versions" now names the latest release and `main`.
+
## [0.1.3] — 2026-09-11
### Fixed
diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md
index c3b9fd9..8e58633 100644
--- a/CONTRIBUTING.md
+++ b/CONTRIBUTING.md
@@ -117,7 +117,7 @@ exit country, and what you
got. Product counts differ by country and by URL, so a bare "worked for me" is
not reproducible.
-Do not add anything that submits the registration form. This project
+Do not add anything that submits a registration or login form. This project
deliberately never does, and a captcha token proved valid by creating a real
account is not a result worth having.
diff --git a/SECURITY.md b/SECURITY.md
index 711081a..36631a4 100644
--- a/SECURITY.md
+++ b/SECURITY.md
@@ -55,7 +55,7 @@ Not because these do not matter, but because they belong somewhere else:
- **Bypassing Amazon's bot protection.** This scraper drives an ordinary
browser and passes challenges the way a browser does. Anything about how
- Akamai or reCAPTCHA behave is not a vulnerability in this repository.
+ AWS WAF or Amazon's own captcha behave is not a vulnerability in this repository.
- **The scraper stopped working.** Amazon changing its markup is expected —
file it as a normal issue, there is a template for exactly that.
- **Anything about 2Captcha's services** — the solver API, the Scraping Browser
@@ -79,9 +79,9 @@ Not because these do not matter, but because they belong somewhere else:
## Supported versions
-`main` only. This project has no releases or version tags; fixes land on `main`
-and you update by pulling. If you are running an old clone, update before
-reporting.
+The latest release and `main`. Fixes land on `main` first and ship in the
+next tagged release (see the Releases page). If you are running an old clone,
+update before reporting.
## If you have leaked a key
diff --git a/TROUBLESHOOTING.md b/TROUBLESHOOTING.md
index 34f55d1..87bda31 100644
--- a/TROUBLESHOOTING.md
+++ b/TROUBLESHOOTING.md
@@ -11,9 +11,8 @@ answer this in one look.
| What the dump shows | Cause |
|---|---|
| A challenge or "verify you are human" page | Bot management. Use a browser engine (not a plain HTTP fetch), a residential IP, or a remote browser via `--cdp-endpoint`. |
-| A real page, prices visible, still 0 rows | The JSON-LD path found nothing and the CSS fallback did not match. Check that product links still match `-item-.aspx`. |
+| A real page, prices visible, still 0 rows | The page kind's data-attribute anchor (`[data-asin]`, `[id^="p13n-asin-index-"]`, `[data-hook="reviewContainer"]`) found nothing and the `/dp/{ASIN}` URL-pattern fallback did not match either. See "How it parses" in the README. |
| A real page in a different language, prices like `125 €` | Fine — that parses. If rows are still 0, it is not the locale. |
-| A near-empty page | The hub URL. `/shopping/kids/items.aspx` has zero products in its JSON-LD; use a filtered category URL. |
## A local Selenium session will not start
diff --git a/diff_runs.py b/diff_runs.py
index 04d4080..2a69fbb 100644
--- a/diff_runs.py
+++ b/diff_runs.py
@@ -8,15 +8,15 @@
monitoring and assortment tracking, but that nothing in this repo actually
computed.
- python3 diff_runs.py --old girls_clothing.2026-09-01.json \\
- --new girls_clothing.2026-09-07.json
+ python3 diff_runs.py --old headphones.2026-09-01.json \\
+ --new headphones.2026-09-07.json
Typical use is a scheduled re-run of one of the four scraper engines, kept
under a dated filename, diffed against the previous one:
- python3 playwright_scraper.py --url "$URL" --out "girls_$(date +%F)"
- python3 diff_runs.py --old "girls_$(ls -t girls_*.json | sed -n 2p)" \\
- --new "girls_$(date +%F).json" --out diff.json
+ python3 playwright_scraper.py --url "$URL" --out "headphones_$(date +%F)"
+ python3 diff_runs.py --old "headphones_$(ls -t headphones_*.json | sed -n 2p)" \\
+ --new "headphones_$(date +%F).json" --out diff.json
Four buckets, each keyed on sku:
@@ -26,9 +26,9 @@
changed — sku present in both, with a different price,
original_price, discount_pct, currency or in_stock
source_changed — sku present in both with a different price, but also a
- different price_source: one run got the DOM-corrected
- figure and the other the raw JSON-LD one, so the two are
- not comparable on price. Reported separately because this
+ different price_source: the two runs read the price from
+ different DOM nodes (offscreen / split / detail), so the
+ two are not comparable on price. Reported separately because this
says something about our own two snapshots, not about the
site — and --fail-on-change deliberately ignores it.