Upload your documents. DocuQuery reads them, indexes them, and answers questions with the exact page and passage the answer came from — no more skimming a 40-page report for one number.
Live demo · Quick start · Features · API reference · Roadmap
Most "chat with your docs" demos are a thin wrapper around one LLM call. DocuQuery is the plumbing that makes that call trustworthy: it parses your files properly, chunks and embeds them, stores the vectors in Postgres, retrieves the passages that actually answer the question, and hands the model only that — then shows you exactly where the answer came from.
- 📂 Drop in a file — PDF, DOCX, TXT, or Markdown, up to 50MB
- ⚡ It gets indexed automatically — parsed, cleaned, chunked, embedded, and stored in seconds
- 💬 Ask in plain language — no query syntax, no filters to configure
- 📎 Get an answer with receipts — every claim links back to the filename, page, and passage
- 🔀 Pick how strict you want it — answer strictly from your documents, from general knowledge, or both
| Step | What happens |
|---|---|
| 1. Upload | Drag in a PDF, DOCX, TXT, or MD file. It's validated, saved, and marked uploaded. |
| 2. Index | In the background: parsed → cleaned → split into passages → embedded into 384-dim vectors → stored in pgvector with an HNSW index. Status flips to indexed. |
| 3. Ask | Type a question in plain language. Pick a mode: answer from your documents only, from the model's general knowledge, or both. |
| 4. Answer | The response streams back token by token, then a citation appears — the exact document, page, and passage the claim came from. |
| Feature | What it means for you |
|---|---|
| 📎 Answers with receipts | Every claim is linked to a filename, page, and passage — click through and verify it yourself |
| 🔀 Three chat modes | DocuQuery (documents only), LLM (general knowledge), Hybrid (documents + reasoning + optional live web search) — switch mid-conversation |
| 🔴 Real-time streaming | Token-by-token responses over SSE, with lossless Markdown — lists, tables, and code blocks render exactly as the model wrote them |
| 📁 A library you can scan | A dense table of every upload with live indexing status and file size |
| 🧮 Real vector search | PostgreSQL + pgvector with HNSW indexing — cosine similarity, not keyword matching |
| 🔌 Bring your own LLM | Groq, OpenAI, Anthropic, Gemini, or a local Ollama model — swap providers via one config value |
| 🔐 Real authentication | Signed HTTP-only session cookies, PBKDF2-SHA256 password hashing, protected routes |
| 🛡️ Production hardening | Per-IP rate limiting on upload and chat, security headers on every response |
| 🎙️ Voice input | Ask out loud via the Web Speech API |
| ⌨️ Built for keyboard users | Ctrl+K for a new chat, full keyboard navigation through menus and the mode selector |
| 🌙 Dark-first design | A deliberate, token-based design system — not default framework styling |
|
Frontend
|
Backend
|
Node.js 20+ · Python 3.13+ · Docker Desktop · a free Groq API key
cd frontend
npm install
npm run devOpen http://localhost:3000.
cd backend
cp .env.example .env
# set LLM_API_KEY and DATABASE_URL in .env
docker compose up --build
docker compose exec api python -m alembic upgrade headAPI runs at http://localhost:8000 · interactive docs at /docs when DEBUG=true.
Prefer running the backend without Docker?
cd backend
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # macOS/Linux
pip install -r requirements.txt
cp .env.example .env
python -m alembic upgrade head
uvicorn app.main:app --reloadEnvironment variables you'll actually need to touch
DATABASE_URL=postgresql+asyncpg://postgres:postgres@localhost:5432/docuquery
LLM_PROVIDER=groq
LLM_MODEL=qwen/qwen3.8-27b
LLM_API_KEY=<your-groq-api-key>
# Optional — enables live web search in Hybrid mode
TAVILY_API_KEY=<your-tavily-api-key>The full list — chunking, retrieval thresholds, rate limits, auth cookie settings — lives in
backend/.env.example with inline comments.
Every request to POST /chat carries a mode:
| Mode | Retrieval | Sources shown |
|---|---|---|
docuquery |
Searches your indexed documents only | Document, page, and passage citations |
llm |
Skips document retrieval entirely | None |
hybrid |
Documents, plus live web search via Tavily when configured | Document citations + linked web sources |
// POST /chat
{ "message": "What are the key findings in the Q3 report?", "mode": "docuquery" }// Response
{
"answer": "Revenue grew 23% YoY to $4.2M...",
"citations": [{ "filename": "Q3_Report_2024.pdf", "page": 7, "chunk_index": 12 }],
"model": "qwen/qwen3.8-27b"
}| Method | Endpoint | Description |
|---|---|---|
GET |
/health |
Liveness + DB connectivity check |
POST |
/upload |
Upload a document — triggers the ingestion pipeline |
GET |
/documents |
List all documents with processing status |
GET |
/documents/{id} |
Get a single document |
GET |
/documents/{id}/chunks |
Inspect a document's indexed chunks |
DELETE |
/documents/{id} |
Delete a document and all its chunks |
POST |
/retrieve |
Vector similarity search — chunks + context, no LLM call |
POST |
/chat |
Full RAG pipeline — streams an answer with citations |
POST |
/auth/signup |
Create an account, establish a session |
POST |
/auth/login |
Authenticate, establish a session |
GET |
/auth/me |
Return the authenticated user |
POST |
/auth/logout |
Clear the session cookie |
Full interactive docs at /docs when the backend runs with DEBUG=true.
How a document becomes searchable
POST /upload
│
▼
Validate file (type + size) ──► saved to disk, status: UPLOADED
│
▼ [background task]
Parse (PyMuPDF · python-docx · plain text, heading-aware for Markdown)
│
▼ status: PROCESSING
Clean text → Chunk (recursive splitter, 800 tokens / 120 overlap)
│
▼
Embed (BAAI/bge-small-en-v1.5, ONNX INT8, 384 dims, batched)
│
▼
Store (PostgreSQL + pgvector, HNSW index, cosine similarity)
│
▼
status: INDEXED
How a question becomes a cited answer
POST /chat { "message": "..." }
│
▼
Embed the query → Vector search (top-k chunks) → Score, dedupe, rank
│
▼
Build context (token-budgeted) → Construct prompt (versioned system prompt)
│
▼
LLM generation (Groq / OpenAI / Anthropic / Gemini / Ollama)
│
▼
Stream response over SSE: tokens → citations → [DONE]
The streaming client strips only the SSE framing space after data: — whitespace, paragraph
breaks, lists, tables, and code blocks from the model arrive intact.
- Document ingestion, vector retrieval, and multi-provider LLM streaming
- Signed session authentication and protected routes
- Rate limiting and security headers
- Hybrid retrieval with BM25 + reciprocal rank fusion
- Cross-encoder reranking
- Role-based access control
- Agentic RAG (multi-agent orchestration)
- Multimodal RAG (OCR, images, tables)
cd backend && pytest -q
cd frontend && npx tsc --noEmit && npm run lint && npm run buildThe backend suite covers parsing, chunking, retrieval scoring, citations, prompt construction, provider abstraction, streaming, password hashing, and session validation.
- Fork the repo and create a branch:
git checkout -b feature/your-feature - Commit using Conventional Commits:
feat: add retrieval endpoint - Run
npm run build(frontend) andpytest(backend) — zero failures before opening a PR - Open a PR against
main
Distributed under the MIT License.



