Skip to content
anjali-patel21Public

About

Pixel-native visual document retrieval for commercial credit review using SigLIP embeddings and FAISS.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

PixelRAG — Commercial Credit Review Assistant (POC)

Pixel-native visual document retrieval for commercial credit review experiments.

What is PixelRAG?

PixelRAG (pixel-native RAG) indexes documents as images instead of extracted text. Each PDF page is rendered to pixels, optionally split into overlapping tiles, and embedded with a vision–language model. Queries are embedded in the same space and matched with vector search (FAISS).

This preserves visual structure that text-only RAG often loses:

  • table rows and columns
  • charts and figures
  • form layouts and label–value pairs
  • headings, sections, and scanned pages

Technical inspiration: Pixel-Native RAG — MarkTechPost.

Why commercial credit review?

A commercial loan review bundle typically includes financial statements, loan agreements, appraisals, and credit memos. Critical facts often live in tables, covenants, valuation charts, and visual risk indicators. This POC tests whether visual retrieval can surface the right page evidence for questions like covenant limits, FY2025 EBITDA, or collateral value — before any agentic credit analysis.

Current architecture

PDF Documents
      ↓
Render Pages
      ↓
Page Images
      ↓
Optional Image Tiling
      ↓
Vision Embedding Model
      ↓
FAISS Vector Index
      ↓
User Query
      ↓
Query Embedding
      ↓
Similarity Search
      ↓
Top Matching Page Images
Module Role
src/ingestion/ PDF → page images; optional overlapping tiles
src/embeddings/ Pluggable visual embedder (siglip / clip / hash)
src/indexing/ FAISS + JSON metadata
src/retrieval/ Query embed → top-K hits (+ optional contact sheet)

V1 stops at question → visual retrieval → document/page evidence. No Vision-LLM answering, OCR hybrid, or agentic workflow yet.

Project setup

Requirements: Python 3.11+

python -m venv .venv

# Windows
.venv\Scripts\activate

# macOS / Linux
source .venv/bin/activate

pip install -r requirements.txt
cp .env.example .env

First SigLIP run downloads model weights from Hugging Face (~hundreds of MB).

Hardware notes

Backend Typical needs
siglip (default: google/siglip-base-patch16-224) CPU OK; GPU optional and faster
clip Similar to SigLIP
Upstream Qwen3-VL-Embedding-2B (not wired in V1) ~8 GB VRAM recommended

Set DEVICE=cpu or DEVICE=cuda in .env to override auto-detect.

How to add documents

  1. Place PDFs in data/documents/ (or the path in DOCUMENT_DIR).
  2. Rebuild the index (below).

Generate the synthetic banking pack used in the test scenario:

python scripts/create_sample_docs.py

This writes:

  • financial_statement_2025.pdf — FY2024/FY2025 tables (Revenue, EBITDA, Debt)
  • loan_agreement.pdf — Debt/EBITDA covenant ≤ 3.5x
  • property_appraisal.pdf — $18.4M valuation + declining chart + risk indicator
  • credit_memo.pdf — “leverage remains within acceptable covenant limits” (intentional contradiction)

How to build the index

python scripts/build_index.py

Example output:

Found 4 documents
Rendered 10 pages
Generated N image chunks
Embedding chunks... done
Created FAISS index with N vectors
Index saved successfully

Artifacts:

  • page/tile images → data/rendered_pages/
  • FAISS index + metadata → data/indexes/

How to query the index

python scripts/query.py "What is the maximum allowed Debt to EBITDA ratio?"

Optional contact sheet of retrieved images:

python scripts/query.py "What is the current property valuation?" --contact-sheet

Example result shape:

Top Results

1. loan_agreement.pdf
   Page: 2
   Tile: 0
   Score: 0.3124
   Image: data/rendered_pages/loan_agreement/...

Suggested test queries:

  1. What is the maximum allowed Debt to EBITDA ratio?
  2. What was the borrower’s FY2025 EBITDA?
  3. How much debt does the borrower have in FY2025?
  4. What is the current property valuation?
  5. Which page contains the Debt to EBITDA covenant?
  6. Show me information related to borrower leverage.
  7. Find evidence related to collateral valuation.
  8. What information shows deterioration in the borrower’s financial condition?

Configuration

See .env.example. Important knobs:

Variable Purpose
DOCUMENT_DIR / RENDERED_IMAGE_DIR / INDEX_DIR Data paths
VISUAL_EMBEDDING_BACKEND siglip (default), clip, or hash (smoke test)
VISUAL_EMBEDDING_MODEL HF model id
ENABLE_TILING / TILE_WIDTH / TILE_HEIGHT / TILE_OVERLAP Image chunking
TOP_K Default retrieval depth
RENDER_DPI PDF rasterization DPI

Embedding interface (swap models without touching retrieval/index code):

class VisualEmbedder:
    def embed_image(self, image): ...
    def embed_query(self, query: str): ...

Current limitations

  • Dense visual embeddings only (no OCR / BM25 hybrid).
  • SigLIP/CLIP are caption-style dual encoders: strong on layout/topic, weaker on fine print than document VLMs.
  • No Vision-LLM answer generation yet (images are retrieval evidence only).
  • No ratio calculation or contradiction detection in V1.
  • Exact FAISS (IndexFlatIP); fine for POC scale, not a large corpus.
  • Synthetic PDFs are clean digital renders, not noisy scans.

Planned experiments

  1. Measure hit-rate on the eight credit queries (page-level Recall@k).
  2. Compare full-page vs tiled indexing (ENABLE_TILING).
  3. Swap SigLIP ↔ CLIP ↔ stronger document embedding models.
  4. Add OCR hybrid + reciprocal rank fusion (as in the PixelRAG tutorial).
  5. Pass top tiles to a Vision LLM for grounded answers.
  6. Detect the leverage contradiction (memo “in covenant” vs 4.2x vs 3.5x max).
  7. Evaluate on real scanned credit packages.

License

POC / experimental — not for production credit decisions.

About

Pixel-native visual document retrieval for commercial credit review using SigLIP embeddings and FAISS.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages