Skip to content

Repository files navigation

hallucite logo

hallucite

Finds fabricated ("hallucinated") references in academic paper PDF files. Each reference is checked against academic databases (the offline DBLP mirror, then CrossRef, DOI resolution, arXiv, OpenAlex and Semantic Scholar); references that no database can confirm are escalated to an interactive LLM triage step, which writes a report for human review.

Three stages: extract and verify use no LLM (verification queries the online databases unless --offline restricts it to the offline DBLP mirror); triage is the only step that uses an LLM, which can be a cloud or a local model. See PLAN.md for the design and architecture.

One repo serves as the runnable project (the mise tasks below) and one shared plugin tree for Claude Code and Codex CLI. The Claude metadata is under .claude-plugin/; the Codex metadata lives under .codex-plugin/ and .agents/plugins/marketplace.json. The bundled skill in skills/hallucite/ drives the same scripts for both tools.

Setup (once)

Run from this directory.

Dependency Needed for Install
mise provisions Python and uv see mise docs
pdftotext (poppler) reference extraction shells out to it brew install poppler
sqlite3 build-dblp checks the database it just built ships with macOS
Playwright + Chromium fetch-dblp-dump -- --daily only; needs a display uv pip install playwright && playwright install chromium
mise install               # provision Python + uv (auto-venv)
mise run fetch-dblp-dump   # download the newest monthly dblp snapshot (~1 GB)
mise run build-dblp        # build ~/hallucite/dblp.db from it (~5 min)

The pipeline itself has no Python dependencies: extraction, parsing, verification and the DBLP ingest are all standard library, and pdftotext is the only outside program it calls.

The offline DBLP database lives at ~/hallucite/dblp.db, outside this repo, which keeps the 3.8 GB file out of git. Set $HALLUCITE_DBLP to store it somewhere else.

API keys

Put them in .env.local in the repo root, one NAME=value per line. The file is gitignored, and both mise and skills/hallucite/scripts/run.sh read it from the tree they run in, so the tasks below and a plugin installed from a local checkout see the same values. A variable already set in the environment wins.

A plugin installed from the marketplace is a different tree and never sees your clone's .env.local. Export the values in the shell that starts Claude Code or Codex CLI, or pass --s2-api-key and --openalex-api-key.

S2_API_KEY=s2k-...        # Semantic Scholar: https://www.semanticscholar.org/product/api
OPENALEX_API_KEY=...      # OpenAlex, optional: https://openalex.org/
CROSSREF_MAILTO=you@...   # CrossRef, optional: a contact address, not a key

Semantic Scholar is not asked without a key. Anonymous callers share a small quota, and a rate-limited lookup leaves a reference degraded rather than cleanly negative, which moves the boundary between verified and "needs triage" between otherwise identical runs. OpenAlex is asked either way and meters the day rather than the second: a hundred searches without a key, a thousand under a free one, reset at midnight UTC. A paper's residue fits the first; a corpus needs the second. CrossRef wants no key. A caller that gives a contact address is routed into a pool allowing three requests at a time where an anonymous one gets a single request, so CROSSREF_MAILTO buys throughput and nothing else. Keeping it in .env.local is what keeps an address off every command line and out of the repository.

Run the audit (Stages 1+2, no LLM)

mise run audit -- <pdf-file-or-dir>            # required: a PDF file, or a directory of PDF files
mise run audit -- <pdf-file-or-dir> [options]  # everything after the target is forwarded as-is

Writes out/<paper_id>.json (every reference plus per-database verification) and out/summary.json (status counts plus the DBLP build date). Options: --dblp PATH, --out DIR, --mailto EMAIL (defaults to $CROSSREF_MAILTO), --s2-api-key KEY, --openalex-api-key KEY, --offline (no network; the offline DBLP mirror stays live), --disable-dbs LIST (comma-separated), --no-verify, --no-candidates (skip the CrossRef lookup that attaches candidate records; implied by --offline), --rate-limit-retries N, --retry-degraded N (re-check what a backend failure left degraded; default 1, 0 disables) and --retry-delay SECONDS (default 5). The DBLP path defaults to $HALLUCITE_DBLP (else ~/hallucite/dblp.db) and the output dir to out. A reference the backends miss is re-verified once with its line-break hyphens removed before it reaches triage. The DBLP backend checks the cited title and authors against every record sharing that title, and where several match it reports the published record the citation locates -- the one whose DOI, page range or volume it prints, and failing those the one whose year it prints. A reference needs triage when its db_verification.status is anything other than verified (not_found, mismatch, or unparsed). Re-running into the same --out is idempotent (triage_verdicts.json accumulates by paper_id:number).

Read the warnings the run ends with, because the audit exits 0 either way. A paper that yielded 0 references was not checked at all, usually a bibliography layout the extractor does not read. An entry number the bibliography prints that no reference carries is a reference that never reached verification, and nothing downstream can report one that never arrived. The per-backend failure tally says how many references came back degraded: their not_found is a weaker claim than a clean negative, and --retry-degraded re-checks them.

Triage the residue (Stage 3, an interactive LLM agent)

mise exec -- python skills/hallucite/scripts/triage.py worklist --out out          # add --pending to skip done
mise exec -- python skills/hallucite/scripts/triage.py worklist --paper <id> --out out  # one paper's slice
mise exec -- python skills/hallucite/scripts/triage.py status --out out             # per-paper done / pending

Stage 3 reads the per-paper JSON the audit has already written, so it can run on finished papers while the audit is still processing the rest. There is no need to wait for the whole corpus. Verdicts accumulate, and worklist --pending surfaces only references not yet recorded. Each worklist entry carries what the audit knew about the reference: which backends matched it and which were never asked, CrossRef's closest records, DBLP's record for the cited title or its nearest title under the same authors, and what a cited DOI or arXiv id resolves to. To fan triage out, hand each worker its own worklist --paper <id> slice (exact id match) instead of the shared worklist, so a worker can't grab the wrong paper (e.g. paper6 vs paper66); record locks the verdicts file, so concurrent workers don't lose each other's verdicts.

Hand the worklist to an interactive LLM agent such as Claude Code or Codex CLI ("triage the unverified references in out"), or use the installed plugin (below). The agent classifies each reference title-first: a partial-match is a real, locatable publication with the cited title but a slipped metadata field (a citation error); a title that matches no real publication is likely-hallucinated, not a partial-match, even when a different paper by the same authors exists. Categories: real-published, real-grey-literature, real-preprint-or-unpublished, partial-match, likely-hallucinated, unclear. The agent records verdicts with structured fabrication signals, then assembles the reports:

mise exec -- python skills/hallucite/scripts/triage.py record <paper_id> <number> <category> "<finding>" \
  --signals '{"title_match":"no","authors_match":"yes","venue_match":"no","doi_status":"none"}' --out out
mise exec -- python skills/hallucite/scripts/triage.py report --out out

record enforces the title-first rule via --signals: partial-match needs title_match=yes (plus a matched_title) or na; likely-hallucinated needs title_match=no. report writes to out/reports/: reference-check-<paper>.md (per paper), potential-hallucinations.md (corpus rollup for review: a severity table, then a Desk-reject candidates section listing references whose cited title matches no real publication, compounded by a fabricated author set, venue, or DOI), and verify-<paper>.md (a manual-check sheet for each flagged paper, with a per-reference verdict line, the matched title, the signals, and one-click Scholar/Google/DOI/arXiv links). Triage is the slow step that calls an LLM; do one paper at a time unless you ask for the whole corpus.

Updating the offline DBLP database

Recent papers cite recent work, so an out-of-date database produces false "not found" results. The audit checks the database's age at run time and prints a warning when ~/hallucite/dblp.db is more than 30 days old. Rebuild it with mise run build-dblp, which builds to a scratch file and swaps it in only after checking that the result is mirror-sized and that accented author names survived the ingest.

dblp publishes two dumps. A monthly snapshot goes out through Schloss Dagstuhl's DROPS, one release per month under its own DOI (10.4230/dblp.xml.2026-09-01), CC0, served by an ordinary web server with an MD5 beside it. dblp.org also publishes a daily dump, fronted by an Anubis proof-of-work bot check that no plain HTTP client answers: curl receives the challenge page instead of the dump, and an ingest reads that as zero publications.

mise run fetch-dblp-dump takes the newest snapshot, checks it against the published MD5, and refuses to install bytes the release disagrees with or that are not a gzip. It needs no browser and no package. Add -- --daily for dblp.org's daily dump, which is fresher by up to a month and drives a real browser that answers the challenge with its own JS engine; that path has to run headed, since Anubis refuses a headless browser outright, so on a machine without a display you download the dump elsewhere and copy it over. Either way, point the build at the file:

mise run fetch-dblp-dump   # -> ~/hallucite/dblp.xml.gz, where build-dblp looks for it
mise run build-dblp        # set DBLP_XML_GZ to build from a dump kept elsewhere

Why the ingest is ours

DBLP writes Latin-1 letters as the named entities its DTD declares (M&aacute;rcio Ribeiro), and an ingest that cannot resolve them drops the author entirely: journals/tse/SoaresRGAS23 kept 3 of its 5, books/sp/WohlinRHOR00 3 of its 6. DBLP's own records are complete; the loss happened on ingest, and it is unsound in exactly one direction -- the mirror then reports an author mismatch for references that are cited correctly, 154 of them across a 2065-reference corpus.

build_dblp.py resolves the entities in the byte stream, as numeric character references so the dump's own ISO-8859-1 declaration cannot turn them back into mojibake, and build-dblp refuses to swap in a database whose authors carry no diacritic at all. The audit reads the same property off whatever mirror it is handed and stops holding an absence against a citation when the mirror cannot support one. Checked against 5000 records of a mirror built by the previous Rust ingest, all 5000 agree on title, year, type, electronic edition and the full author list.

Install as a plugin

claude plugin marketplace add se-uhd/hallucite      # GitHub, or a local clone path
claude plugin install hallucite@hallucite

For Codex CLI, add the marketplace and install the plugin:

codex plugin marketplace add se-uhd/hallucite
codex plugin list --marketplace hallucite
codex plugin add hallucite@hallucite

To update later:

codex plugin marketplace upgrade hallucite
codex plugin add hallucite@hallucite

For local development, you can register a local checkout instead:

codex plugin marketplace add /path/to/hallucite
codex plugin add hallucite@hallucite

The local path install uses the plugins/hallucite -> .. compatibility shim and may copy the current working tree into Codex's plugin cache, including ignored local directories. Use a clean checkout when testing local installs.

Then in any session, ask for it in words: "check the references in <dir> for hallucinations". Claude Code also lists the bundled skill as hallucite:hallucite. The skill (skills/hallucite/SKILL.md) resolves the bundled skills/hallucite/scripts/run.sh from a Claude plugin install, a Codex repo-local skill shim, a direct repo clone, the Claude Code plugin cache, or the Codex plugin cache. That wrapper is the single entry point (check-env | audit | triage | lint | python). Installed plugins do not need mise: run.sh finds a Python 3.10 or newer with sqlite3 -- searching PATH and the places a plugin's non-interactive shell tends to miss -- and fails loud with a HALLUCITE_BOOTSTRAP_FAILED: line rather than running half-configured. There is nothing to install: the pipeline is standard library only. Set $HALLUCITE_PYTHON to pin an interpreter. run.sh check-env reports whether the environment is ready, including a warning when pdftotext is missing.

You still build the offline DBLP database once, and without a clone there are no mise tasks to do it with. Get dblp.xml.gz with fetch_dblp_dump.py, which downloads the newest monthly snapshot and verifies its checksum, then run the ingest through the wrapper:

RUN=<plugin>/skills/hallucite/scripts/run.sh
"$RUN" python <plugin>/skills/hallucite/scripts/build_dblp.py ~/hallucite/dblp.xml.gz \
  --out ~/hallucite/dblp.db

mise run build-dblp also checks the result before installing it, which this path does not, so confirm it yourself: over a million publications, and author names carrying a diacritic.

Tests and linting

python skills/hallucite/scripts/tests/run_smoke.py

Measuring a change to extraction or verification, all offline (measure/ in the same directory):

python skills/hallucite/scripts/verification_results.py compare ~/hallucite/verification-results-offline.json --impl verifier
python skills/hallucite/scripts/measure/fabricated_citations.py score ~/hallucite/fabricated-citations.json
python skills/hallucite/scripts/measure/mutations.py

The fabricated-citation set has to come back unchanged from any change to the author or title rules. It is three sets in one file: 250 real records with titles of three or more words and 250 with two-word titles, and 90 whose DBLP row carries an et al. marker, each cited correctly and then with known fabrications. Correct citations confirm; one, two or three invented names appended confirm 0. Score the built file, never rebuild it, because a rebuild on a newer dump samples different records. The offline results file is what a code change is diffed against; a change that moves no verdict replays with no status move.

run_smoke.py is the suite CI runs on every push; its docstring lists what each tier covers. It needs pdftotext for the extraction tiers and nothing else. Markdown lint runs as a separate CI step; locally, mise run lint-md (below).

The repo's Markdown is checked with a vendored PyMarkdown (synced from se-uhd/pymarkdown-skill; self-contained under skills/hallucite/scripts/, no pip install):

mise run lint-md            # check every tracked Markdown file
MD_FIX=1 mise run lint-md   # auto-fix in place

Contributing

Commit messages follow Conventional Commits (type(scope): summary, for example fix(extract): ...); keep the release version out of the message and record it with a v* git tag instead. Run mise run lint-md and the smoke tests (python skills/hallucite/scripts/tests/run_smoke.py) before a release.

About

Find fabricated ("hallucinated") references in academic paper PDF files.

Resources

Stars

15 stars

Watchers

1 watching

Forks

Contributors

Languages