Reliable delivery of Instagram fraud reports, from real authenticated sessions.
It reports what it requested. It cannot know what Instagram did with them.
That sentence is the design, not a disclaimer. Instagram renders a success message optimistically to reporters it does not trust, so there is no oracle: no response, no toast, and no DOM state can distinguish "filed" from "accepted and discarded". Every claim this tool makes is therefore about its own request — what it sent, from which exit, under which identity, recorded before it left — and never about the target account. If you need "the account was removed", no tool can tell you that, and one that claims to is lying.
This section is first because everything after it is worthless without it.
| Agreed definition of done | State |
|---|---|
1. doctor passes on at least one channel from a clean exit |
not done |
| 2. A report submitted, network response and state transition both in the ledger | not done |
| 3. Killed mid-submit, resumed, nothing resubmitted | offline coverage only |
| 4. No credential in any log, artifact, or checkpoint (CI-enforced) | done |
| 5. Selector drift yields a DOM diff naming expected vs actual | done |
| 6. Every anticipated failure mode has a mitigation or a written acceptance | done |
Item 6's acceptances are the ones stated inline here: Instagram's optimistic success rendering (the opening section), and the absence of an oracle for "accepted and discarded".
No report has ever been filed by this tool against a live Instagram. The test suite is green (1119 tests) and it is green with every channel unable to reach Instagram. That sentence is the single most important thing on this page.
Items 1–3 need two things that do not exist yet:
- A residential proxy exit. There is no
proxies.txt— the pool is empty and no file is tracked. Open-proxy harvesting is deliberately unsupported: those addresses are pre-scorched by every scraper on the internet, and a session bound to one is dead before the first request lands. A report filed from your home IP is also the one correlation that gets the reporting account flagged, which is the opposite of the tool's purpose. - Your consent to file a real report from a real account against a real target you control.
The container is written and unproven. Dockerfile and docker-compose.yml
are committed and the compose file parses, but the image has not been built and
nothing has been run from it. The browser channel is the only implemented
channel, and it has never been pointed at live Instagram from anywhere. Treat the
container section below as a design that has not met its first run.
Until both exist, treat the channels as unverified and the terminal vocabulary as a recording format rather than a result.
insta_report/api.py is complete as a classifier and empty as an endpoint.
The mobile/web identity ladder, the request construction, the dispatch boundary
and the response classification are all implemented and tested. The endpoint
itself is not, because nothing about it has ever been observed answering.
ApiEndpoint refuses to be constructed unverified, and insta_report/cli.py does
not import the channel at all. Two independent locks, both with tests that fail if
removed. A fallback you have not verified is a black hole — that is the mistake
which caused this rewrite, and the code refuses to repeat it.
Requires Python 3.11+.
git clone https://github.com/reblox01/InstaReport.git
cd InstaReport
python -m venv .venv
.venv\Scripts\pip install -e ".[dev]" # Windows
.venv\Scripts\python -m playwright install chromiumCopy the example and keep it out of version control — it is gitignored for a
reason documented at the top of .gitignore.
cp config.example.toml config.tomlPoint [api] and [browser] at real values, and put your session cookie in the
environment rather than in the file:
$env:IG_SESSIONID_ALPHA = "<paste>"sessionid_env names the variable holding the cookie, so no credential is ever
written to disk in the config. To get the cookie: log in on instagram.com, open
devtools → Application → Cookies → sessionid.
You also need each reporting account's own numeric user_id. It is not a secret,
so unlike the session it lives in the file. Without it the probe refuses to run
rather than address the reporting route with a literal {user_id} in the path —
a malformed path's 404 cannot be told apart from a missing route, so it is not
output.
insta-report doctor --probe-target <an account you control>Always start here. doctor gates the runner by exit code and run refuses to
start if no channel passes. Skip it with --no-doctor and you have made the first
run of this tool an unverified production run.
insta-report run --targets targets.txt # deliver
insta-report run --dry-run # print the plan, send nothing
insta-report status --run <id> # what may have been reported
insta-report targets --targets targets.txt # confusables, duplicates, self-reports
insta-report anchors --check dump.json # selector driftThe observability probe is not a subcommand — it is a module entry point, and it is read-only by construction (it asserts every request it issues is a read, and refuses to build a URL with an unsubstituted placeholder):
python -m insta_report.probe --config config.toml --username <a handle you control>run writes a durable checkpoint before every submit, so an interrupted run
resumes without double-reporting. status lists the targets that were dispatched
and never resolved — the ones that may have been reported.
docker compose build
docker compose run --rm reporter doctor --no-live # exits 1. That is correct.
docker compose run --rm reporter run --targets /targets/accounts.txtThree things are worth knowing before you use it, and two of them are traps.
The image exists for Chromium, not for isolation. The pin fixes the browser
revision, not the base image — playwright==1.63.0 bundles a browsers.json
naming exact Chromium revisions, so playwright install chromium under that pin
fetches revision 1243 whichever base you build on. The base is
mcr.microsoft.com/playwright/python:v1.63.0-noble for the part the pin cannot
do: playwright install --with-deps resolves the browser's shared libraries from
an Ubuntu package list, and python:3.11-slim is Debian. So it is a 2.5 GB base
that cannot fail at build time over a small one that can, because the browser
channel has never been run against live Instagram at all and "which build failed"
is not a question worth leaving open.
It does not give you a different IP. A container shares its host's network namespace, so every lease still egresses from your address and the own-address guard will dismiss it. Docker earns its place on a second host — a VPS — where that host's address is genuinely not yours, and for the reproducible browser. On this laptop it is packaging, not privacy. What actually moves the egress is a residential proxy, and nothing in this repository changes that.
The sessionid is supplied at run time, from outside the repository.
docker-compose.yml reads env_file, defaulting to ../insta-report.env — one
level up, so a secrets file cannot land in the checkout by accident. Override
it on a VPS:
INSTA_REPORT_ENV_FILE=/run/secrets/insta-report.env docker compose run --rm reporter doctorThe file holds the bare cookie value, not the cookie:
IG_SESSIONID_ALPHA=<the value of the sessionid cookie>
It is never an ARG or an ENV in the image. That is not a style preference: a
build argument is permanent, sitting in the image, the build cache, and
docker history, and this repository is public, so a pushed image would be a
published cookie. The image does declare IG_SESSIONID_ALPHA= — present and
empty — so that a container started with no credential refuses to run rather than
authenticating with a blank one.
Stated plainly, because the alternative is overclaiming: the value is visible
to docker inspect on the running container, since a process cannot read its own
environment otherwise. The protection is that it never reaches git, the image, or
the build cache — not that it is hidden from the machine running it.
INSTA_REPORT_DATA_DIR is the one config value a container may override.
Your config's data_dir is a Windows path that does not exist in the container,
and the tool refuses a data directory inside the work tree, so it would be
refused in /app as well. Compose sets it to /data, which is a bind mount onto
./artifacts — a bind mount rather than a named volume because
docker compose down -v would destroy a named volume, and that volume is the
record of what was attempted against real accounts. Create artifacts/ and
targets/ on the host first, owned by uid 1000: Docker silently creates a
missing bind-mount source as a directory, and a missing proxies.txt then
becomes a directory that is unreadable as a file, with a permission error
standing in for "the file is not there".
The offline suite runs in its own container, which is the stronger place to prove the offline claim because it is a clean machine:
docker compose run --rm --profile verify verifyThis is the tool's real output. Every report attempt ends in exactly one.
| state | meaning | stops_ladder |
counts_against_budget |
|---|---|---|---|
SUBMITTED_ACKED |
dispatched, response and DOM confirm it | yes | yes |
SUBMITTED_UNCONFIRMED |
dispatched, response readable but not a success | yes | yes |
UNKNOWN |
dispatched, response uninterpretable | yes | yes |
NOT_REPORTABLE |
target gone or unresolvable | yes | no |
CHANNEL_FAILED |
this channel could not attempt it | no | no |
QUARANTINED |
not attempted; the exit is out of service | yes | no |
needs_human_review is true for exactly two states — SUBMITTED_UNCONFIRMED and
UNKNOWN — because those are the two where the tool cannot tell you what
Instagram did. That property, not a heuristic in your shell script, is how you
find the rows to check by hand.
The dispatch boundary is the load-bearing line. Before it, a failure is
recoverable — try the next channel, try the next exit. After it, nothing is
retried and nothing falls through, because a report that may have been filed
cannot be safely filed again. CHANNEL_FAILED continues the ladder precisely
because it means the server rejected the request or it never left: a 4xx is a
refusal to process, so no report exists.
insta_report/
outcomes.py terminal vocabulary + the golden corpus that pins it
errors.py Transient / ChannelFail / Fatal, each scoped REPORT/LEASE/RUN
checkpoint.py JSONL, fsync per record, intent-before-dispatch
runner.py ladder orchestration, dispatch boundary, per-channel health
browser.py Playwright channel (the primary path)
api.py API channel — classifier only, disarmed
proxies.py lease pool, health, egress-IP assertion
accounts.py budgets, LRU rotation, quarantine 2^n
pacing.py monotonic, budget-derived cadence
anchors.py the selector oracle + drift reporting
artifacts.py evidence bundles, redacted on write
narrative.py report text rendering
doctor.py the gate
probe.py observability probe
cli.py command surface
Dockerfile Playwright image, pinned to the driver version
docker-compose.yml bind-mounted config and artifacts, env_file for the secret
.dockerignore active rules matter: the build context is sent whole
.venv\Scripts\python -m pytest # 1119 tests, offline
.venv\Scripts\python -m pytest -m "not static" # skip pyflakes + mypy
.venv\Scripts\python -m pytest -m "not browser" # skip the Chromium testsThe suite never touches the network — and that is now enforced, not promised. An
autouse fixture in tests/conftest.py fails any test that opens a non-loopback
socket, so a test that reaches the internet fails loudly instead of passing
because the network happened to be up. Loopback is allowed on purpose: asyncio
holds a self-pipe per event loop and on Windows that is a real socket on
127.0.0.1. A test that genuinely needs the network must ask with
@pytest.mark.network_access; a test asserts that nothing is so marked, so
adding one is a visible act that forces this section to be corrected.
That guard exists because the suite was not offline. Adding the own-address
check below put a live request to api.ipify.org on the path that builds the
exit pool, and roughly twenty run tests started depending on a third party's
uptime without anyone deciding to. They passed in CI and would have failed on a
plane. Two layers do reach the network, and both are invoked by hand: doctor
and the probe module.
Six things are enforced as tests rather than conventions, because each one found a real defect:
- pyflakes found two test functions whose names shadowed each other, so one had never run in the life of the suite.
- mypy found a declaration that lied (
_anchors: AnchorSetdefaulting toNone), and following the lie found five unguarded reads that would have been reported to an operator as Instagram rejected the session. - A credential scan of everything git tracks, plus a
git add -ACI test, because.gitignorecannot stop a developer pasting a real cookie into a fixture. Secrets are covered on the wire and in artifacts, not only in logs. - A pattern for a sessionid value carrying no cookie name. Every other
pattern anchors on a name —
sessionid=,apikey:,scheme://user:pass@— and a test fixture has no name, becauseregister_secret()takes the bare value. So the shape that actually occurs in practice was the one shape nothing matched, and this repository shipped a fixture built from the operator's real account id for the whole life of the API work without the gate noticing. No secret was in it, only an identifier that should not have been published — which is precisely the case a credential scanner is blind to by construction. The pattern is structural (ds_user_iddigit run,%3Aseparators, length and case mixing), souser_id = "61214264580"still does not fire. - The offline socket guard itself. A guard nothing tests is a convention again, so the gate is exercised both ways: a non-loopback connect must be refused, and a loopback connect must be allowed. The first of those failed during development for a reason worth recording — the guard is what caught the ipify regression, and nothing else did.
- This test count. It had already gone stale twice — 1088, then 1105, while the suite quietly grew past both — so the number is now read out of pytest's own collection and compared. A figure in a README is a claim about the suite, and a claim nobody checks is how the other claims in this file got made.
The container adds four more, in tests/test_repository_hygiene.py: the pinned
dependency set in the image must equal the one pyproject declares, no ARG or
ENV may carry a credential, the compose file may not define the sessionid
inline, and its default secrets path must resolve outside the checkout. Each was
verified by breaking it. That last one was wrong on the first attempt — it
matched a substring, so commenting a rule out satisfied it, which is the same
trap this repository's own .gitignore documents.
Every lease is checked against the address this machine reports for itself, and a lease that egresses from it is refused. That is the one correlation the tool exists to avoid: a report filed from the address Instagram already associates with you identifies the reporting account, and no amount of session hygiene compensates for it.
operator's own IP ──┐
├── canonical form ── equal? ── yes ──> leased
lease's egress IP ──┘ └── no ──> bind
Left unset in [proxies], the address is observed once at startup with a direct,
non-proxied request, and a pool that could not observe it still builds but refuses
to lease — the refusal names both remedies, and an unreachable IP-echo service
deserves one line rather than a traceback out of a constructor. Set own_ip when
that cannot be inferred correctly: behind CGNAT, behind a corporate egress that
rewrites source addresses, or when the echo service is unreachable from your
network. A value there that is not a valid address is refused, not ignored —
leave the key out rather than in with something unchecked.
An address that turns out to be the operator's own is dismissed, not
quarantined, and the distinction is the whole design. Quarantine is for addresses
that might come back; this one is a configuration error that will not fix itself.
And it is filtered out of the candidate list rather than only refused at bind
time, so the answer stays the same for acquire, for the fallback pass that
runs under load, and for an operator reading status.
Two invariants in the probe are refused rather than checked after the fact, because both defects produced output that looked like evidence: a probe may not issue a non-read method, and a probe URL may not carry an unsubstituted placeholder. A 404 nobody can interpret is not a result.
igban.py and get_id.py are the original single-file script, left in place and
unmaintained. They are what this rewrite exists to replace, and the reason is
specific rather than stylistic — igban.py decides success from a status code:
if response.status_code == 200:
...
# If response isn't JSON (e.g. HTML confirmation), just say success
print(f"[+] User {user_id} reported successfully.")It reports success most confidently in the case the new vocabulary calls
UNKNOWN — a 200 whose body it could not read.
The other half is a dispatch-boundary violation in the opposite direction. On any non-200 it immediately re-POSTs the identical payload to a second URL:
# Fallback: Try the primary web endpoint again just in case
resp_alt = session.post(url_alt, headers=headers, data=data)A non-200 includes 429 and 5xx — precisely the responses where the request may still have been processed. So an ambiguous first response is answered with a guaranteed second dispatch, and the tool can file the same report twice. Neither behaviour is fixable by improving the parsing: the response is not the evidence. The six states above exist so that "I sent this and I cannot tell you what happened" is a thing the tool can say out loud, instead of a gap that gets papered over with a second request.
MIT, see LICENSE. For reporting accounts that violate Instagram's terms. Use
where you are the affected party; the cost of a false report is borne by whoever
the report names.