Skip to content

Repository files navigation

🌱 PlantMind Copilot

The Unified Asset & Operations Brain for Industrial Intelligence

Turning 7–12 disconnected plant document systems into one queryable, self-updating knowledge graph.

Python FastAPI Neo4j ChromaDB Docker License

Quick Start β€’ Architecture β€’ Features β€’ API Reference β€’ Testing


πŸ“– The Problem

A 2024 McKinsey survey of asset-intensive industries found that engineers and technicians spend 35% of their working hours searching for information, clarifying instructions, or recreating documents that already exist somewhere in the organization. A NASSCOM-EY study found the typical large industrial plant runs 7 to 12 disconnected document systems β€” P&IDs in one place, maintenance work orders in another, SOPs in a third, inspection records in a fourth, regulatory filings scattered across email archives.

The cost isn't inconvenience. It's compliance blind spots, root-cause analysis that takes days instead of minutes, and β€” as India's industrial sector faces a wave of retirements over the next decade β€” decades of undocumented operational knowledge walking out the door with the people who held it.

PlantMind is built to close that gap: ingest any industrial document format, extract and link the entities inside it, and make the result queryable in seconds β€” with citations, confidence scores, and zero silent failures.


πŸ“ Project Overview

PlantMind Copilot is an AI-powered Industrial Knowledge Intelligence platform designed to eliminate information silos across industrial plants.

Instead of storing SOPs, inspection reports, maintenance logs, incident reports, work orders, regulations, and emails in isolated systems, PlantMind transforms them into a unified knowledge graph combined with semantic search.

By combining Hybrid Retrieval-Augmented Generation (Hybrid RAG), Neo4j Knowledge Graphs, ChromaDB Vector Search, and AI Agents, PlantMind enables engineers to ask complex operational questions and receive explainable, citation-backed answers within seconds.

The goal is to reduce time spent searching documentation, improve compliance readiness, accelerate root-cause analysis, and preserve critical operational knowledge before it is lost.


πŸ—οΈ Architecture

flowchart TD
    subgraph Input["πŸ“₯ Heterogeneous Document Input"]
        A1[PDF - native & scanned]
        A2[DOCX]
        A3[XLSX]
        A4[EML]
        A5[TXT / MD / CSV]
    end

    subgraph Ingestion["βš™οΈ Universal Ingestion Pipeline"]
        B1[Format Router]
        B2[Parser: pypdf / python-docx / openpyxl / email-stdlib]
        B3[Content-Hash Dedup β€” SHA-256]
        B4[Paragraph-Aware Chunker]
    end

    subgraph Extraction["🧠 Entity Extraction Layer"]
        C1[LLM Structured Extraction<br/>via OpenRouter]
        C2[Pydantic Schema Validation<br/>+ Retry on Malformed JSON]
        C3[Entity Normalizer<br/>Pump/Compressor/Valve/Vessel β†’ Canonical Tag]
    end

    subgraph Storage["πŸ’Ύ Dual Storage Layer"]
        D1[(ChromaDB<br/>Vector Embeddings)]
        D2[(Neo4j 5<br/>Knowledge Graph<br/>Multi-Tenant Compound Keys)]
    end

    subgraph Intelligence["πŸ€– Reasoning Layer"]
        E1[Hybrid RAG Copilot<br/>Vector + Multi-Hop Graph]
        E2[Compliance Agent]
        E3[RCA Agent]
        E4[Lessons Learned Agent]
        E5[Knowledge Decay Engine]
        E6[Compound Risk Finder]
    end

    subgraph Output["πŸ“€ Delivery"]
        F1[Streaming Chat β€” SSE]
        F2[Interactive Graph β€” vis.js]
        F3[Downloadable Audit Reports]
    end

    Input --> Ingestion
    Ingestion --> Extraction
    Extraction --> Storage
    Storage --> Intelligence
    Intelligence --> Output

    style Input fill:#1a1500,stroke:#E8A33D,color:#E7E4D9
    style Ingestion fill:#0e2a1a,stroke:#6FBF73,color:#E7E4D9
    style Extraction fill:#2a1a0e,stroke:#E8A33D,color:#E7E4D9
    style Storage fill:#1a1a2a,stroke:#4A7CE0,color:#E7E4D9
    style Intelligence fill:#2a0e0e,stroke:#D9645C,color:#E7E4D9
    style Output fill:#1a1500,stroke:#E8A33D,color:#E7E4D9
Loading

πŸŽ₯ Demo Video

Watch PlantMind Copilot in action:

πŸ“Ή Project Demo:
https://drive.google.com/file/d/1LV7amU769ucO4Zid3F3y2ue8jUFX2AxO/view?usp=sharing

Why Graph + Vector, not just Vector

Plain vector search retrieves similar text. It cannot answer "what regulations apply to the pump mentioned in yesterday's incident report?" β€” that answer is assembled from two documents that share no vocabulary at all. PlantMind's Copilot finds the incident report via semantic similarity, extracts its equipment tag, then traverses the graph to the regulation node connected two hops away. This hybrid retrieval path is the system's core technical differentiator.

sequenceDiagram
    participant U as User Query
    participant V as ChromaDB (Vector)
    participant G as Neo4j (Graph)
    participant L as LLM (OpenRouter)

    U->>V: "What regulations apply to P-204's incident?"
    V-->>U: Top match: Incident_Report_P204.txt
    U->>G: Multi-hop expand from Incident_Report_P204
    G-->>U: 2 hops away: OISD-STD-113 (via Equipment P-204)
    U->>L: Context = [Incident Report + OISD-STD-113]
    L-->>U: Cited, confidence-scored answer
Loading

✨ Core Features

Feature What It Does Why It Matters
🧠 Hybrid RAG Copilot Merges ChromaDB semantic search with Neo4j multi-hop graph traversal Answers cross-document questions pure vector RAG structurally cannot
πŸ”₯ Knowledge Decay Heatmap Tracks time-since-last-documented per equipment node, colors it Healthy / Stale / Critical against OISD-STD-113-anchored thresholds Makes the "retiring engineer knowledge cliff" visible and actionable, not abstract
πŸ•΅οΈ Compound Risk Finder Graph query surfacing document pairs that share equipment or zone context but were never cross-referenced Catches the exact failure pattern behind real industrial incidents β€” signal that existed but was never connected
πŸ›‘οΈ Compliance Agent Audits ingested regulations against procedures and inspection records, generates a downloadable Markdown evidence package Turns a multi-hour manual audit-prep task into a one-click export
πŸ”§ Maintenance & RCA Agent Fuses work orders, inspection findings, and incident records into root-cause hypotheses and predictive maintenance recommendations Connects dots no single team member sees across fragmented systems
πŸ’‘ Lessons Learned Agent Scans incident/near-miss history for systemic patterns, drafts proactive warnings Surfaces recurring failure modes before they repeat
πŸ“„ Universal Parser Native text-layer PDFs, scanned PDFs (graceful degradation), DOCX, XLSX, EML, CSV, TXT, MD Handles the real, messy heterogeneity of actual plant documentation
πŸ”’ Content-Hash Deduplication SHA-256 of content + tenant scope as the document key Re-ingesting the same file never fragments the graph
🏒 True Multi-Tenancy Every graph node keyed on (id, plant_id) compound key Two plants using identical equipment naming (P-204) never cross-contaminate
⚑ Streaming Responses Server-Sent Events for the Copilot No frozen spinner β€” tokens render live

⏱️ Time-to-Answer Impact (illustrative β€” validate against your own corpus before citing as measured)

Task Traditional Approach PlantMind
Find all documents referencing a piece of equipment 15–25 min across disconnected systems Seconds, via Copilot
Check equipment compliance against a regulation 45–90 min cross-referencing 3+ systems Seconds, via Compliance Agent
Identify cross-document risk patterns Rarely done systematically at all Seconds, via Compound Risk Finder
Assemble a compliance evidence package for audit Hours of manual document assembly One click, via Audit Export

πŸ› οΈ Tech Stack

Layer Technology Role
Backend API FastAPI + Uvicorn Async REST + SSE streaming
Graph Database Neo4j 5 Multi-tenant knowledge graph, multi-hop traversal
Vector Database ChromaDB Semantic document retrieval
LLM Orchestration OpenRouter (OpenAI-compatible client) Model-agnostic β€” Claude, Llama, or any listed model
Structured Extraction Pydantic Schema-validated entity extraction with retry
Document Parsing pypdf, python-docx, openpyxl, stdlib email Native multi-format ingestion, zero heavy ML dependencies
Frontend Vanilla JS + HTML/CSS + vis.js Zero build step, fully portable, interactive graph rendering
Deployment Docker Compose One-command reproducible environment
CI GitHub Actions Lint (ruff) + automated test suite on every push

πŸš€ Quick Start

Prerequisites

1. Clone the repository

git clone https://github.com/AnmollCodes/PlantMind-Copilot.git
cd PlantMind-Copilot

2. Configure environment

Create a .env file in the project root:

OPENROUTER_API_KEY=sk-or-v1-your-key-here
OPENROUTER_MODEL=anthropic/claude-3.5-sonnet
NEO4J_PASSWORD=choose-a-password
OCR_PROVIDER=none

Any model listed at openrouter.ai/models works β€” swap OPENROUTER_MODEL without touching code.

3. Build and run

docker compose up --build

Neo4j has a health check β€” the API container automatically waits for the graph database to be ready before starting.

4. Verify it's alive

curl http://localhost:8000/health

5. Open the app

Navigate to http://localhost:8000 in your browser.

6. Try the included demo corpus

The demo_docs/ folder contains a realistic, interconnected set of industrial documents spanning every supported format (SOP, Inspection Report, Work Order, Incident Report, Shift Log, Regulation, Maintenance Record, plus a spreadsheet and an email). Upload them through the UI, watch the knowledge graph build live, then ask the Copilot:

"What abnormal readings were reported for pump P-204, and what does OISD-STD-113 require for its inspection?"

You'll get a cited answer pulling from across the email, the spreadsheet, and the regulation document simultaneously.


πŸ“‘ API Reference

Method Endpoint Purpose
GET /health Liveness check β€” always returns 200, even with no credentials configured
GET /health/detailed Confirms Neo4j and ChromaDB connectivity
POST /ingest Upload and process a document (multipart form: file, plant_id, doc_type)
POST /query Ask the RAG Copilot a question (question, plant_id, mode: hybrid|vector)
POST /query/stream Same as above, streamed via SSE
POST /agents/{agent_key} Run rca, compliance, or lessons agent for a plant
GET /agents/compliance/export/{plant_id} Download a Markdown compliance audit package
GET /decay/{plant_id} Knowledge decay status for every equipment node
GET /compound-risk/{plant_id} Cross-document risk pattern findings
GET /graph/{plant_id} / /graph/{plant_id}/stats Full graph data / node-count summary
GET /docs/list/{plant_id} List all ingested documents for a plant

Every endpoint is scoped by plant_id β€” full multi-tenant isolation, enforced at the graph query level via compound node keys, not just application-layer filtering.


πŸ§ͺ Testing & Quality

python -m venv venv
source venv/bin/activate      # Windows: venv\Scripts\Activate.ps1
pip install -r requirements.txt

pytest tests/test_core_logic.py -v
ruff check app/

The test suite covers entity normalization (including boundary cases for equipment prefix families), chunking, schema validation, agent empty-corpus short-circuits, knowledge-decay boundary conditions, audit export formatting, content-hash deduplication, compound-risk query deduplication, and multi-tenant isolation under shared equipment tags across different plants β€” the last of which was found and fixed as a real compound-key MERGE bug during development, not assumed correct.

CI runs linting and the full suite on every push via GitHub Actions (.github/workflows/).


πŸ“ Project Structure

PlantMind-Copilot/
β”œβ”€β”€ app/
β”‚   β”œβ”€β”€ main.py            # FastAPI app, all endpoints, lazy-init clients
β”‚   β”œβ”€β”€ ingestion.py        # Format router + 7 document parsers
β”‚   β”œβ”€β”€ extraction.py       # LLM entity extraction + Pydantic schema + normalizer
β”‚   β”œβ”€β”€ graph.py             # Neo4j queries: linking, multi-hop, compound risk, decay
β”‚   β”œβ”€β”€ agents.py            # RCA, Compliance, Lessons Learned agents
β”‚   β”œβ”€β”€ decay.py             # Pure-function knowledge decay classification
β”‚   └── audit_export.py      # Compliance report Markdown formatting
β”œβ”€β”€ tests/
β”‚   └── test_core_logic.py   # Full test suite β€” logic-level, no live DB required
β”œβ”€β”€ demo_docs/                # Realistic multi-format demo corpus
β”œβ”€β”€ docker-compose.yml         # Neo4j + API, one-command deployment
β”œβ”€β”€ Dockerfile
└── requirements.txt

πŸ—ΊοΈ Roadmap

Honestly scoped, not oversold β€” these are deliberately not in the current build:

  • P&ID / engineering drawing parsing β€” requires a computer-vision layer over drawing files; the current parser correctly handles text-bearing documents but a P&ID's information lives in its image layer, not its text layer.
  • Voice-to-knowledge ingestion for field technicians in PPE gloves.
  • Live IoT/SCADA telemetry fusion for real-time predictive maintenance, beyond the current document-based analysis.
  • OCR provider integration (Mistral OCR / LlamaParse) for scanned documents β€” the extension point exists in ingestion.py, provider-agnostic by design, not yet wired to a live vendor.

🚧 Project Status

PlantMind Copilot is actively under development.

The current version includes:

  • βœ… Hybrid RAG (Vector + Graph Retrieval)
  • βœ… Multi-format document ingestion
  • βœ… Knowledge Graph visualization
  • βœ… Compliance Agent
  • βœ… Root Cause Analysis Agent
  • βœ… Lessons Learned Agent
  • βœ… Knowledge Decay Analysis
  • βœ… Compound Risk Detection
  • βœ… Multi-tenant architecture

Upcoming improvements include:

  • πŸ”„ OCR integration for scanned engineering documents
  • πŸ”„ P&ID diagram understanding
  • πŸ”„ Voice-based field engineer assistant
  • πŸ”„ IoT & SCADA real-time integration
  • πŸ”„ Advanced Agentic Workflows
  • πŸ”„ Better UI/UX and production deployment

The project is continuously evolving with new capabilities and improvements.


πŸ“œ License

MIT β€” see LICENSE.


Built to eliminate industrial knowledge silos. 🏭

About

Industrial Knowledge Intelligence platform built for complex operational environments. Merges vector search with Neo4j graph traversal for audit-grade answers and automated compliance reporting.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages