Turning 7β12 disconnected plant document systems into one queryable, self-updating knowledge graph.
Quick Start β’ Architecture β’ Features β’ API Reference β’ Testing
A 2024 McKinsey survey of asset-intensive industries found that engineers and technicians spend 35% of their working hours searching for information, clarifying instructions, or recreating documents that already exist somewhere in the organization. A NASSCOM-EY study found the typical large industrial plant runs 7 to 12 disconnected document systems β P&IDs in one place, maintenance work orders in another, SOPs in a third, inspection records in a fourth, regulatory filings scattered across email archives.
The cost isn't inconvenience. It's compliance blind spots, root-cause analysis that takes days instead of minutes, and β as India's industrial sector faces a wave of retirements over the next decade β decades of undocumented operational knowledge walking out the door with the people who held it.
PlantMind is built to close that gap: ingest any industrial document format, extract and link the entities inside it, and make the result queryable in seconds β with citations, confidence scores, and zero silent failures.
PlantMind Copilot is an AI-powered Industrial Knowledge Intelligence platform designed to eliminate information silos across industrial plants.
Instead of storing SOPs, inspection reports, maintenance logs, incident reports, work orders, regulations, and emails in isolated systems, PlantMind transforms them into a unified knowledge graph combined with semantic search.
By combining Hybrid Retrieval-Augmented Generation (Hybrid RAG), Neo4j Knowledge Graphs, ChromaDB Vector Search, and AI Agents, PlantMind enables engineers to ask complex operational questions and receive explainable, citation-backed answers within seconds.
The goal is to reduce time spent searching documentation, improve compliance readiness, accelerate root-cause analysis, and preserve critical operational knowledge before it is lost.
flowchart TD
subgraph Input["π₯ Heterogeneous Document Input"]
A1[PDF - native & scanned]
A2[DOCX]
A3[XLSX]
A4[EML]
A5[TXT / MD / CSV]
end
subgraph Ingestion["βοΈ Universal Ingestion Pipeline"]
B1[Format Router]
B2[Parser: pypdf / python-docx / openpyxl / email-stdlib]
B3[Content-Hash Dedup β SHA-256]
B4[Paragraph-Aware Chunker]
end
subgraph Extraction["π§ Entity Extraction Layer"]
C1[LLM Structured Extraction<br/>via OpenRouter]
C2[Pydantic Schema Validation<br/>+ Retry on Malformed JSON]
C3[Entity Normalizer<br/>Pump/Compressor/Valve/Vessel β Canonical Tag]
end
subgraph Storage["πΎ Dual Storage Layer"]
D1[(ChromaDB<br/>Vector Embeddings)]
D2[(Neo4j 5<br/>Knowledge Graph<br/>Multi-Tenant Compound Keys)]
end
subgraph Intelligence["π€ Reasoning Layer"]
E1[Hybrid RAG Copilot<br/>Vector + Multi-Hop Graph]
E2[Compliance Agent]
E3[RCA Agent]
E4[Lessons Learned Agent]
E5[Knowledge Decay Engine]
E6[Compound Risk Finder]
end
subgraph Output["π€ Delivery"]
F1[Streaming Chat β SSE]
F2[Interactive Graph β vis.js]
F3[Downloadable Audit Reports]
end
Input --> Ingestion
Ingestion --> Extraction
Extraction --> Storage
Storage --> Intelligence
Intelligence --> Output
style Input fill:#1a1500,stroke:#E8A33D,color:#E7E4D9
style Ingestion fill:#0e2a1a,stroke:#6FBF73,color:#E7E4D9
style Extraction fill:#2a1a0e,stroke:#E8A33D,color:#E7E4D9
style Storage fill:#1a1a2a,stroke:#4A7CE0,color:#E7E4D9
style Intelligence fill:#2a0e0e,stroke:#D9645C,color:#E7E4D9
style Output fill:#1a1500,stroke:#E8A33D,color:#E7E4D9
Watch PlantMind Copilot in action:
πΉ Project Demo:
https://drive.google.com/file/d/1LV7amU769ucO4Zid3F3y2ue8jUFX2AxO/view?usp=sharing
Plain vector search retrieves similar text. It cannot answer "what regulations apply to the pump mentioned in yesterday's incident report?" β that answer is assembled from two documents that share no vocabulary at all. PlantMind's Copilot finds the incident report via semantic similarity, extracts its equipment tag, then traverses the graph to the regulation node connected two hops away. This hybrid retrieval path is the system's core technical differentiator.
sequenceDiagram
participant U as User Query
participant V as ChromaDB (Vector)
participant G as Neo4j (Graph)
participant L as LLM (OpenRouter)
U->>V: "What regulations apply to P-204's incident?"
V-->>U: Top match: Incident_Report_P204.txt
U->>G: Multi-hop expand from Incident_Report_P204
G-->>U: 2 hops away: OISD-STD-113 (via Equipment P-204)
U->>L: Context = [Incident Report + OISD-STD-113]
L-->>U: Cited, confidence-scored answer
| Feature | What It Does | Why It Matters |
|---|---|---|
| π§ Hybrid RAG Copilot | Merges ChromaDB semantic search with Neo4j multi-hop graph traversal | Answers cross-document questions pure vector RAG structurally cannot |
| π₯ Knowledge Decay Heatmap | Tracks time-since-last-documented per equipment node, colors it Healthy / Stale / Critical against OISD-STD-113-anchored thresholds | Makes the "retiring engineer knowledge cliff" visible and actionable, not abstract |
| π΅οΈ Compound Risk Finder | Graph query surfacing document pairs that share equipment or zone context but were never cross-referenced | Catches the exact failure pattern behind real industrial incidents β signal that existed but was never connected |
| π‘οΈ Compliance Agent | Audits ingested regulations against procedures and inspection records, generates a downloadable Markdown evidence package | Turns a multi-hour manual audit-prep task into a one-click export |
| π§ Maintenance & RCA Agent | Fuses work orders, inspection findings, and incident records into root-cause hypotheses and predictive maintenance recommendations | Connects dots no single team member sees across fragmented systems |
| π‘ Lessons Learned Agent | Scans incident/near-miss history for systemic patterns, drafts proactive warnings | Surfaces recurring failure modes before they repeat |
| π Universal Parser | Native text-layer PDFs, scanned PDFs (graceful degradation), DOCX, XLSX, EML, CSV, TXT, MD | Handles the real, messy heterogeneity of actual plant documentation |
| π Content-Hash Deduplication | SHA-256 of content + tenant scope as the document key | Re-ingesting the same file never fragments the graph |
| π’ True Multi-Tenancy | Every graph node keyed on (id, plant_id) compound key |
Two plants using identical equipment naming (P-204) never cross-contaminate |
| β‘ Streaming Responses | Server-Sent Events for the Copilot | No frozen spinner β tokens render live |
β±οΈ Time-to-Answer Impact (illustrative β validate against your own corpus before citing as measured)
| Task | Traditional Approach | PlantMind |
|---|---|---|
| Find all documents referencing a piece of equipment | 15β25 min across disconnected systems | Seconds, via Copilot |
| Check equipment compliance against a regulation | 45β90 min cross-referencing 3+ systems | Seconds, via Compliance Agent |
| Identify cross-document risk patterns | Rarely done systematically at all | Seconds, via Compound Risk Finder |
| Assemble a compliance evidence package for audit | Hours of manual document assembly | One click, via Audit Export |
| Layer | Technology | Role |
|---|---|---|
| Backend API | FastAPI + Uvicorn | Async REST + SSE streaming |
| Graph Database | Neo4j 5 | Multi-tenant knowledge graph, multi-hop traversal |
| Vector Database | ChromaDB | Semantic document retrieval |
| LLM Orchestration | OpenRouter (OpenAI-compatible client) | Model-agnostic β Claude, Llama, or any listed model |
| Structured Extraction | Pydantic | Schema-validated entity extraction with retry |
| Document Parsing | pypdf, python-docx, openpyxl, stdlib email |
Native multi-format ingestion, zero heavy ML dependencies |
| Frontend | Vanilla JS + HTML/CSS + vis.js | Zero build step, fully portable, interactive graph rendering |
| Deployment | Docker Compose | One-command reproducible environment |
| CI | GitHub Actions | Lint (ruff) + automated test suite on every push |
- Docker Desktop installed and running
- An OpenRouter API key (free-tier models available)
git clone https://github.com/AnmollCodes/PlantMind-Copilot.git
cd PlantMind-CopilotCreate a .env file in the project root:
OPENROUTER_API_KEY=sk-or-v1-your-key-here
OPENROUTER_MODEL=anthropic/claude-3.5-sonnet
NEO4J_PASSWORD=choose-a-password
OCR_PROVIDER=noneAny model listed at openrouter.ai/models works β swap
OPENROUTER_MODELwithout touching code.
docker compose up --buildNeo4j has a health check β the API container automatically waits for the graph database to be ready before starting.
curl http://localhost:8000/healthNavigate to http://localhost:8000 in your browser.
The demo_docs/ folder contains a realistic, interconnected set of industrial documents spanning every supported format (SOP, Inspection Report, Work Order, Incident Report, Shift Log, Regulation, Maintenance Record, plus a spreadsheet and an email). Upload them through the UI, watch the knowledge graph build live, then ask the Copilot:
"What abnormal readings were reported for pump P-204, and what does OISD-STD-113 require for its inspection?"
You'll get a cited answer pulling from across the email, the spreadsheet, and the regulation document simultaneously.
| Method | Endpoint | Purpose |
|---|---|---|
GET |
/health |
Liveness check β always returns 200, even with no credentials configured |
GET |
/health/detailed |
Confirms Neo4j and ChromaDB connectivity |
POST |
/ingest |
Upload and process a document (multipart form: file, plant_id, doc_type) |
POST |
/query |
Ask the RAG Copilot a question (question, plant_id, mode: hybrid|vector) |
POST |
/query/stream |
Same as above, streamed via SSE |
POST |
/agents/{agent_key} |
Run rca, compliance, or lessons agent for a plant |
GET |
/agents/compliance/export/{plant_id} |
Download a Markdown compliance audit package |
GET |
/decay/{plant_id} |
Knowledge decay status for every equipment node |
GET |
/compound-risk/{plant_id} |
Cross-document risk pattern findings |
GET |
/graph/{plant_id} / /graph/{plant_id}/stats |
Full graph data / node-count summary |
GET |
/docs/list/{plant_id} |
List all ingested documents for a plant |
Every endpoint is scoped by plant_id β full multi-tenant isolation, enforced at the graph query level via compound node keys, not just application-layer filtering.
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\Activate.ps1
pip install -r requirements.txt
pytest tests/test_core_logic.py -v
ruff check app/The test suite covers entity normalization (including boundary cases for equipment prefix families), chunking, schema validation, agent empty-corpus short-circuits, knowledge-decay boundary conditions, audit export formatting, content-hash deduplication, compound-risk query deduplication, and multi-tenant isolation under shared equipment tags across different plants β the last of which was found and fixed as a real compound-key MERGE bug during development, not assumed correct.
CI runs linting and the full suite on every push via GitHub Actions (.github/workflows/).
PlantMind-Copilot/
βββ app/
β βββ main.py # FastAPI app, all endpoints, lazy-init clients
β βββ ingestion.py # Format router + 7 document parsers
β βββ extraction.py # LLM entity extraction + Pydantic schema + normalizer
β βββ graph.py # Neo4j queries: linking, multi-hop, compound risk, decay
β βββ agents.py # RCA, Compliance, Lessons Learned agents
β βββ decay.py # Pure-function knowledge decay classification
β βββ audit_export.py # Compliance report Markdown formatting
βββ tests/
β βββ test_core_logic.py # Full test suite β logic-level, no live DB required
βββ demo_docs/ # Realistic multi-format demo corpus
βββ docker-compose.yml # Neo4j + API, one-command deployment
βββ Dockerfile
βββ requirements.txt
Honestly scoped, not oversold β these are deliberately not in the current build:
- P&ID / engineering drawing parsing β requires a computer-vision layer over drawing files; the current parser correctly handles text-bearing documents but a P&ID's information lives in its image layer, not its text layer.
- Voice-to-knowledge ingestion for field technicians in PPE gloves.
- Live IoT/SCADA telemetry fusion for real-time predictive maintenance, beyond the current document-based analysis.
- OCR provider integration (Mistral OCR / LlamaParse) for scanned documents β the extension point exists in
ingestion.py, provider-agnostic by design, not yet wired to a live vendor.
PlantMind Copilot is actively under development.
The current version includes:
- β Hybrid RAG (Vector + Graph Retrieval)
- β Multi-format document ingestion
- β Knowledge Graph visualization
- β Compliance Agent
- β Root Cause Analysis Agent
- β Lessons Learned Agent
- β Knowledge Decay Analysis
- β Compound Risk Detection
- β Multi-tenant architecture
Upcoming improvements include:
- π OCR integration for scanned engineering documents
- π P&ID diagram understanding
- π Voice-based field engineer assistant
- π IoT & SCADA real-time integration
- π Advanced Agentic Workflows
- π Better UI/UX and production deployment
The project is continuously evolving with new capabilities and improvements.
MIT β see LICENSE.
Built to eliminate industrial knowledge silos. π