Pixel-native visual document retrieval for commercial credit review experiments.
PixelRAG (pixel-native RAG) indexes documents as images instead of extracted text. Each PDF page is rendered to pixels, optionally split into overlapping tiles, and embedded with a vision–language model. Queries are embedded in the same space and matched with vector search (FAISS).
This preserves visual structure that text-only RAG often loses:
- table rows and columns
- charts and figures
- form layouts and label–value pairs
- headings, sections, and scanned pages
Technical inspiration: Pixel-Native RAG — MarkTechPost.
A commercial loan review bundle typically includes financial statements, loan agreements, appraisals, and credit memos. Critical facts often live in tables, covenants, valuation charts, and visual risk indicators. This POC tests whether visual retrieval can surface the right page evidence for questions like covenant limits, FY2025 EBITDA, or collateral value — before any agentic credit analysis.
PDF Documents
↓
Render Pages
↓
Page Images
↓
Optional Image Tiling
↓
Vision Embedding Model
↓
FAISS Vector Index
↓
User Query
↓
Query Embedding
↓
Similarity Search
↓
Top Matching Page Images
| Module | Role |
|---|---|
src/ingestion/ |
PDF → page images; optional overlapping tiles |
src/embeddings/ |
Pluggable visual embedder (siglip / clip / hash) |
src/indexing/ |
FAISS + JSON metadata |
src/retrieval/ |
Query embed → top-K hits (+ optional contact sheet) |
V1 stops at question → visual retrieval → document/page evidence. No Vision-LLM answering, OCR hybrid, or agentic workflow yet.
Requirements: Python 3.11+
python -m venv .venv
# Windows
.venv\Scripts\activate
# macOS / Linux
source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .envFirst SigLIP run downloads model weights from Hugging Face (~hundreds of MB).
| Backend | Typical needs |
|---|---|
siglip (default: google/siglip-base-patch16-224) |
CPU OK; GPU optional and faster |
clip |
Similar to SigLIP |
Upstream Qwen3-VL-Embedding-2B (not wired in V1) |
~8 GB VRAM recommended |
Set DEVICE=cpu or DEVICE=cuda in .env to override auto-detect.
- Place PDFs in
data/documents/(or the path inDOCUMENT_DIR). - Rebuild the index (below).
Generate the synthetic banking pack used in the test scenario:
python scripts/create_sample_docs.pyThis writes:
financial_statement_2025.pdf— FY2024/FY2025 tables (Revenue, EBITDA, Debt)loan_agreement.pdf— Debt/EBITDA covenant ≤ 3.5xproperty_appraisal.pdf— $18.4M valuation + declining chart + risk indicatorcredit_memo.pdf— “leverage remains within acceptable covenant limits” (intentional contradiction)
python scripts/build_index.pyExample output:
Found 4 documents
Rendered 10 pages
Generated N image chunks
Embedding chunks... done
Created FAISS index with N vectors
Index saved successfully
Artifacts:
- page/tile images →
data/rendered_pages/ - FAISS index + metadata →
data/indexes/
python scripts/query.py "What is the maximum allowed Debt to EBITDA ratio?"Optional contact sheet of retrieved images:
python scripts/query.py "What is the current property valuation?" --contact-sheetExample result shape:
Top Results
1. loan_agreement.pdf
Page: 2
Tile: 0
Score: 0.3124
Image: data/rendered_pages/loan_agreement/...
Suggested test queries:
- What is the maximum allowed Debt to EBITDA ratio?
- What was the borrower’s FY2025 EBITDA?
- How much debt does the borrower have in FY2025?
- What is the current property valuation?
- Which page contains the Debt to EBITDA covenant?
- Show me information related to borrower leverage.
- Find evidence related to collateral valuation.
- What information shows deterioration in the borrower’s financial condition?
See .env.example. Important knobs:
| Variable | Purpose |
|---|---|
DOCUMENT_DIR / RENDERED_IMAGE_DIR / INDEX_DIR |
Data paths |
VISUAL_EMBEDDING_BACKEND |
siglip (default), clip, or hash (smoke test) |
VISUAL_EMBEDDING_MODEL |
HF model id |
ENABLE_TILING / TILE_WIDTH / TILE_HEIGHT / TILE_OVERLAP |
Image chunking |
TOP_K |
Default retrieval depth |
RENDER_DPI |
PDF rasterization DPI |
Embedding interface (swap models without touching retrieval/index code):
class VisualEmbedder:
def embed_image(self, image): ...
def embed_query(self, query: str): ...- Dense visual embeddings only (no OCR / BM25 hybrid).
- SigLIP/CLIP are caption-style dual encoders: strong on layout/topic, weaker on fine print than document VLMs.
- No Vision-LLM answer generation yet (images are retrieval evidence only).
- No ratio calculation or contradiction detection in V1.
- Exact FAISS (
IndexFlatIP); fine for POC scale, not a large corpus. - Synthetic PDFs are clean digital renders, not noisy scans.
- Measure hit-rate on the eight credit queries (page-level Recall@k).
- Compare full-page vs tiled indexing (
ENABLE_TILING). - Swap SigLIP ↔ CLIP ↔ stronger document embedding models.
- Add OCR hybrid + reciprocal rank fusion (as in the PixelRAG tutorial).
- Pass top tiles to a Vision LLM for grounded answers.
- Detect the leverage contradiction (memo “in covenant” vs 4.2x vs 3.5x max).
- Evaluate on real scanned credit packages.
POC / experimental — not for production credit decisions.