Skip to content
VampiricCyborgPublic

About

A production-grade, full-stack AI assistant platform for private document fleets — featuring real-time streaming, intelligent document ingestion, vector search, LLM-powered answers, and voice input.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

49 Commits

Folders and files

Repository files navigation

DocuQuery — ask your documents a question and get a cited answer

DocuQuery

Turn a folder of PDFs, DOCX, and text files into a source you can ask questions.

Upload your documents. DocuQuery reads them, indexes them, and answers questions with the exact page and passage the answer came from — no more skimming a 40-page report for one number.

Live demo · Quick start · Features · API reference · Roadmap

Build Status Version Next.js React FastAPI PostgreSQL License PRs Welcome


What DocuQuery actually does

Most "chat with your docs" demos are a thin wrapper around one LLM call. DocuQuery is the plumbing that makes that call trustworthy: it parses your files properly, chunks and embeds them, stores the vectors in Postgres, retrieves the passages that actually answer the question, and hands the model only that — then shows you exactly where the answer came from.

  • 📂 Drop in a file — PDF, DOCX, TXT, or Markdown, up to 50MB
  • ⚡ It gets indexed automatically — parsed, cleaned, chunked, embedded, and stored in seconds
  • 💬 Ask in plain language — no query syntax, no filters to configure
  • 📎 Get an answer with receipts — every claim links back to the filename, page, and passage
  • 🔀 Pick how strict you want it — answer strictly from your documents, from general knowledge, or both

👀 See it work

A question is typed, DocuQuery searches the uploaded documents, streams back an answer, then shows the cited source
One question, start to finish — typed, searched, answered, cited. No cuts.

DocuQuery landing page in light mode

🧭 How it works

Four steps: upload, index, ask, answer
Step What happens
1. Upload Drag in a PDF, DOCX, TXT, or MD file. It's validated, saved, and marked uploaded.
2. Index In the background: parsed → cleaned → split into passages → embedded into 384-dim vectors → stored in pgvector with an HNSW index. Status flips to indexed.
3. Ask Type a question in plain language. Pick a mode: answer from your documents only, from the model's general knowledge, or both.
4. Answer The response streams back token by token, then a citation appears — the exact document, page, and passage the claim came from.

✨ Features

Every answer shows its receipts, three chat modes, and a document library you can scan
Feature What it means for you
📎 Answers with receipts Every claim is linked to a filename, page, and passage — click through and verify it yourself
🔀 Three chat modes DocuQuery (documents only), LLM (general knowledge), Hybrid (documents + reasoning + optional live web search) — switch mid-conversation
🔴 Real-time streaming Token-by-token responses over SSE, with lossless Markdown — lists, tables, and code blocks render exactly as the model wrote them
📁 A library you can scan A dense table of every upload with live indexing status and file size
🧮 Real vector search PostgreSQL + pgvector with HNSW indexing — cosine similarity, not keyword matching
🔌 Bring your own LLM Groq, OpenAI, Anthropic, Gemini, or a local Ollama model — swap providers via one config value
🔐 Real authentication Signed HTTP-only session cookies, PBKDF2-SHA256 password hashing, protected routes
🛡️ Production hardening Per-IP rate limiting on upload and chat, security headers on every response
🎙️ Voice input Ask out loud via the Web Speech API
⌨️ Built for keyboard users Ctrl+K for a new chat, full keyboard navigation through menus and the mode selector
🌙 Dark-first design A deliberate, token-based design system — not default framework styling

🛠️ Tech stack

Frontend

Framework Next.js 16.2 (App Router)
Language TypeScript 5
Styling Tailwind CSS v4
UI primitives Radix UI
Animation Framer Motion
State Zustand 5 + TanStack Query v5
Markdown react-markdown + remark-gfm

Backend

Framework FastAPI 0.115 on Python 3.13
Database PostgreSQL 16 + pgvector
ORM / migrations SQLAlchemy 2 (async) + Alembic
Parsing PyMuPDF (PDF) · python-docx
Embeddings ONNX Runtime — BAAI/bge-small-en-v1.5 (INT8)
LLM providers Groq · OpenAI · Anthropic · Gemini · Ollama
Rate limiting slowapi
Deployment Railway (API) · Vercel (web)

🚀 Quick start

Prerequisites

Node.js 20+ · Python 3.13+ · Docker Desktop · a free Groq API key

1 — Frontend

cd frontend
npm install
npm run dev

Open http://localhost:3000.

2 — Backend (Docker, recommended)

cd backend
cp .env.example .env
# set LLM_API_KEY and DATABASE_URL in .env
docker compose up --build
docker compose exec api python -m alembic upgrade head

API runs at http://localhost:8000 · interactive docs at /docs when DEBUG=true.

Prefer running the backend without Docker?
cd backend
python -m venv .venv
.venv\Scripts\activate        # Windows
# source .venv/bin/activate   # macOS/Linux

pip install -r requirements.txt
cp .env.example .env
python -m alembic upgrade head
uvicorn app.main:app --reload
Environment variables you'll actually need to touch
DATABASE_URL=postgresql+asyncpg://postgres:postgres@localhost:5432/docuquery

LLM_PROVIDER=groq
LLM_MODEL=qwen/qwen3.8-27b
LLM_API_KEY=<your-groq-api-key>

# Optional — enables live web search in Hybrid mode
TAVILY_API_KEY=<your-tavily-api-key>

The full list — chunking, retrieval thresholds, rate limits, auth cookie settings — lives in backend/.env.example with inline comments.


💬 Chat modes

Every request to POST /chat carries a mode:

Mode Retrieval Sources shown
docuquery Searches your indexed documents only Document, page, and passage citations
llm Skips document retrieval entirely None
hybrid Documents, plus live web search via Tavily when configured Document citations + linked web sources
// POST /chat
{ "message": "What are the key findings in the Q3 report?", "mode": "docuquery" }
// Response
{
  "answer": "Revenue grew 23% YoY to $4.2M...",
  "citations": [{ "filename": "Q3_Report_2024.pdf", "page": 7, "chunk_index": 12 }],
  "model": "qwen/qwen3.8-27b"
}

🔌 API reference

Method Endpoint Description
GET /health Liveness + DB connectivity check
POST /upload Upload a document — triggers the ingestion pipeline
GET /documents List all documents with processing status
GET /documents/{id} Get a single document
GET /documents/{id}/chunks Inspect a document's indexed chunks
DELETE /documents/{id} Delete a document and all its chunks
POST /retrieve Vector similarity search — chunks + context, no LLM call
POST /chat Full RAG pipeline — streams an answer with citations
POST /auth/signup Create an account, establish a session
POST /auth/login Authenticate, establish a session
GET /auth/me Return the authenticated user
POST /auth/logout Clear the session cookie

Full interactive docs at /docs when the backend runs with DEBUG=true.

How a document becomes searchable
POST /upload
     │
     ▼
Validate file (type + size) ──► saved to disk, status: UPLOADED
     │
     ▼  [background task]
Parse  (PyMuPDF · python-docx · plain text, heading-aware for Markdown)
     │
     ▼  status: PROCESSING
Clean text  →  Chunk  (recursive splitter, 800 tokens / 120 overlap)
     │
     ▼
Embed  (BAAI/bge-small-en-v1.5, ONNX INT8, 384 dims, batched)
     │
     ▼
Store  (PostgreSQL + pgvector, HNSW index, cosine similarity)
     │
     ▼
status: INDEXED
How a question becomes a cited answer
POST /chat  { "message": "..." }
     │
     ▼
Embed the query  →  Vector search (top-k chunks)  →  Score, dedupe, rank
     │
     ▼
Build context (token-budgeted)  →  Construct prompt (versioned system prompt)
     │
     ▼
LLM generation (Groq / OpenAI / Anthropic / Gemini / Ollama)
     │
     ▼
Stream response over SSE: tokens → citations → [DONE]

The streaming client strips only the SSE framing space after data: — whitespace, paragraph breaks, lists, tables, and code blocks from the model arrive intact.


🗺️ Roadmap

  • Document ingestion, vector retrieval, and multi-provider LLM streaming
  • Signed session authentication and protected routes
  • Rate limiting and security headers
  • Hybrid retrieval with BM25 + reciprocal rank fusion
  • Cross-encoder reranking
  • Role-based access control
  • Agentic RAG (multi-agent orchestration)
  • Multimodal RAG (OCR, images, tables)

🧪 Running tests

cd backend && pytest -q
cd frontend && npx tsc --noEmit && npm run lint && npm run build

The backend suite covers parsing, chunking, retrieval scoring, citations, prompt construction, provider abstraction, streaming, password hashing, and session validation.


🤝 Contributing

  1. Fork the repo and create a branch: git checkout -b feature/your-feature
  2. Commit using Conventional Commits: feat: add retrieval endpoint
  3. Run npm run build (frontend) and pytest (backend) — zero failures before opening a PR
  4. Open a PR against main

📄 License

Distributed under the MIT License.


Built with Next.js, FastAPI, pgvector, and Groq.

About

A production-grade, full-stack AI assistant platform for private document fleets — featuring real-time streaming, intelligent document ingestion, vector search, LLM-powered answers, and voice input.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages