Skip to content

Latest commit

Β 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Cloud//Base β€” Cloud Knowledge Assistant

A production-grade, containerized Retrieval-Augmented Generation (RAG) system engineered to deliver low-latency, interview-grade explanations across core Cloud, DevOps, and Infrastructure domains.

Cloud//Base combines a curated Markdown knowledge base, local HuggingFace embeddings, FAISS semantic retrieval, Gemini, and a Streamlit interface to provide grounded technical answers with transparent source inspection and retrieval telemetry.


πŸ“Έ Interface & Screenshots

1. System Overview & Interface

The main Cloud//Base interface provides system status, supported knowledge domains, retrieval depth, suggested interview topics, and the primary question interface.

Cloud Base UI


2. Grounded Technical Explanation

Cloud//Base transforms retrieved knowledge into a structured technical explanation designed for both learning and interview preparation.

Structured Technical Synthesis


3. Real-Time Latency & Retrieval Telemetry

Each query exposes runtime telemetry including total latency, local vector retrieval latency, Gemini generation time, and the number of retrieved context chunks.

Performance Telemetry Cards


4. Vector Source Inspection

The retrieved knowledge chunks can be inspected directly, including their originating Markdown file, similarity score, and retrieved content.

Grounded Source Chunks


5. Multi-Stage Container Runtime

The application is packaged and executed as a Docker container, providing a reproducible runtime environment for Cloud//Base.

Docker Container Execution


πŸ›οΈ System Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    Knowledge Base (.md)                     β”‚
β”‚  AWS β€’ Docker β€’ Kubernetes β€’ Linux β€’ Terraform β€’ Networking β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β”‚ MarkdownHeaderTextSplitter
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                Local HuggingFace Embeddings                 β”‚
β”‚              all-MiniLM-L6-v2 β€’ CPU β€’ β‚Ή0 Cost              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β”‚ Normalized Vectors
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                     FAISS Vector Store                      β”‚
β”‚            Local Cosine Distance Similarity Index           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β”‚ Top-K Semantic Chunks
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                 Gemini 3.5 Flash Lite LLM                   β”‚
β”‚             Grounded Prompt Synthesis + Fallback            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                 Streamlit Web Interface                     β”‚
β”‚              Dark UI + Real-Time Telemetry                  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ”„ RAG Pipeline

Cloud//Base follows a retrieval-first architecture:

User Question
      β”‚
      β–Ό
Question Embedding
      β”‚
      β–Ό
FAISS Similarity Search
      β”‚
      β–Ό
Top-K Relevant Chunks
      β”‚
      β–Ό
Retrieved Context + User Question
      β”‚
      β–Ό
Grounded Gemini Prompt
      β”‚
      β–Ό
Structured Technical Answer
      β”‚
      β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Ί Retrieved Sources
      β”‚
      └──────────────► Retrieval Telemetry

The system separates the retrieval stage from the generation stage, allowing the retrieved context and source information to be inspected alongside the generated answer.


✨ Key Technical Features

🧠 Retrieval-Augmented Generation

Cloud//Base retrieves relevant information from a curated Cloud, DevOps, and Infrastructure knowledge base before generating an answer.

This allows responses to be grounded in the project's own technical documentation.

⚑ Local Vector Retrieval

FAISS is used as the local vector store for semantic similarity search.

The local retrieval pipeline avoids the need for an externally hosted vector database.

πŸ’° Zero-Cost Embedding Pipeline

Embeddings are generated locally using:

sentence-transformers/all-MiniLM-L6-v2

The embedding model runs locally on CPU without requiring a paid embedding API.

πŸ“š Curated Technical Knowledge

The knowledge base is organized into Markdown files covering:

  • AWS
  • Docker
  • Kubernetes
  • Linux
  • Networking
  • Terraform

🎯 Structured Technical Answers

Responses are organized around technical sections such as:

  1. Executive Definition & Core Value
  2. Architecture & Technical Mechanics
  3. Production Best Practices
  4. Interview Q&A / Key Takeaway

This makes the application useful for both technical learning and interview preparation.

πŸ”Ž Source Grounding

Retrieved chunks can be inspected directly.

The interface exposes information including:

  • Source Markdown file
  • Similarity score
  • Retrieved content
  • Number of retrieved chunks

πŸ“Š Retrieval Telemetry

Each query provides runtime information including:

  • Total latency
  • Local retrieval latency
  • Gemini generation latency
  • Number of context chunks retrieved

🐳 Dockerized Runtime

The application can be packaged into a Docker image and run as a container, providing a consistent runtime environment independent of the local Python installation.


πŸ› οΈ Technology Stack

Layer Technology
Language Python
UI Streamlit
RAG Framework LangChain
LLM Google Gemini
Embeddings HuggingFace Sentence Transformers
Embedding Model all-MiniLM-L6-v2
Vector Store FAISS
Knowledge Format Markdown
Containerization Docker
Version Control Git / GitHub

πŸ“‚ Project Structure

cloud-knowledge-assistant/
β”‚
β”œβ”€β”€ app/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ embeddings.py         # HuggingFace & FAISS index management
β”‚   β”œβ”€β”€ ingestion.py          # Markdown header chunking pipeline
β”‚   β”œβ”€β”€ main.py               # Streamlit UI & telemetry dashboard
β”‚   β”œβ”€β”€ rag.py                # LangChain prompt chaining & Gemini routing
β”‚   └── retrieval.py          # Similarity search helper utilities
β”‚
β”œβ”€β”€ assets/
β”‚   β”œβ”€β”€ 01-landing-ui.png
β”‚   β”œβ”€β”€ 02-query-response.png
β”‚   β”œβ”€β”€ 03-telemetry-metrics.png
β”‚   β”œβ”€β”€ 04-grounded-sources.png
β”‚   └── 05-docker-running.png
β”‚
β”œβ”€β”€ data/
β”‚   └── vectorstore/
β”‚       └── ...
β”‚
β”œβ”€β”€ knowledge/
β”‚   β”œβ”€β”€ aws.md
β”‚   β”œβ”€β”€ docker.md
β”‚   β”œβ”€β”€ kubernetes.md
β”‚   β”œβ”€β”€ linux.md
β”‚   β”œβ”€β”€ networking.md
β”‚   └── terraform.md
β”‚
β”œβ”€β”€ .dockerignore
β”œβ”€β”€ .env.example
β”œβ”€β”€ .gitignore
β”œβ”€β”€ Dockerfile
β”œβ”€β”€ requirements.txt
└── README.md

πŸš€ Quickstart

1. Clone the Repository

git clone https://github.com/<your-username>/cloud-knowledge-assistant.git
cd cloud-knowledge-assistant

2. Create the Environment File

Create a .env file in the project root:

GEMINI_API_KEY="your_google_ai_studio_api_key"
GEMINI_MODEL="gemini-3.5-flash-lite"
EMBEDDING_MODEL_NAME="sentence-transformers/all-MiniLM-L6-v2"
VECTOR_STORE_DIR="data/vectorstore"

Important: Never commit .env or API keys to GitHub.

Use .env.example to document required environment variables without exposing credentials.


🐍 Local Python Execution

Create the Virtual Environment

py -3.12 -m venv venv

Activate the Environment

.\venv\Scripts\Activate.ps1

Verify Python

python --version

Expected:

Python 3.12.x

Install Dependencies

pip install -r requirements.txt

Launch Cloud//Base

streamlit run app/main.py

The application will be available at:

http://localhost:8501

🐳 Docker

Cloud//Base can also be executed inside Docker.

Build the Docker Image

docker build -t cloud-knowledge-assistant:latest .

Run the Container

docker run -d `
  -p 8501:8501 `
  --env-file .env `
  --name cloud-assistant `
  cloud-knowledge-assistant:latest

Access the application at:

http://localhost:8501

Check Running Containers

docker ps

View Container Logs

docker logs cloud-assistant

Stop the Container

docker stop cloud-assistant

Remove the Container

docker rm cloud-assistant

πŸ” Environment & Security

Cloud//Base uses environment variables for API credentials and configuration.

Example:

.env
 β”‚
 β”œβ”€β”€ GEMINI_API_KEY
 β”œβ”€β”€ GEMINI_MODEL
 β”œβ”€β”€ EMBEDDING_MODEL_NAME
 └── VECTOR_STORE_DIR

The real .env file should never be committed to source control.

Recommended .gitignore entries:

.env
venv/
__pycache__/
*.pyc

πŸ§ͺ Example Questions

Cloud//Base can be used to explore technical concepts such as:

Explain ECS vs EKS with an interview answer.

How does chmod 755 work mathematically?

What is the difference between Layer 4 and Layer 7 Load Balancers?

How does Terraform state locking work with DynamoDB?

What is Docker?

What is a Kubernetes Pod?

What is the role of etcd in Kubernetes?

What is a Docker image vs a Docker container?

The retrieved context is shown alongside the generated response so that the grounding of the answer can be inspected.


πŸ“Š Query Telemetry

For each question, Cloud//Base tracks:

Metric Description
Total Latency Total time taken to process the request
Local Retrieval Time spent performing FAISS retrieval
Gemini Generation Time spent generating the final response
Context Chunks Number of knowledge chunks retrieved

Example telemetry from a query:

TOTAL LATENCY       4.965s
LOCAL RETRIEVAL     0.3274s
GEMINI GENERATION   4.638s
CONTEXT CHUNKS      4 retrieved

🎯 Project Objective

The objective of Cloud//Base is to build a practical RAG application where the underlying retrieval and generation pipeline can be understood and inspected rather than treating the LLM as a black box.

The project demonstrates practical experience with:

Python
   β”‚
   β”œβ”€β”€ LangChain
   β”‚
   β”œβ”€β”€ RAG
   β”‚
   β”œβ”€β”€ Embeddings
   β”‚
   β”œβ”€β”€ Vector Search
   β”‚
   β”œβ”€β”€ FAISS
   β”‚
   β”œβ”€β”€ Gemini
   β”‚
   └── Streamlit
          β”‚
          β–Ό
        Docker

πŸ“Œ Current Project Status

Completed

  • Curated Cloud/DevOps knowledge base
  • Markdown document ingestion
  • Markdown header-based chunking
  • Local HuggingFace embeddings
  • FAISS vector store
  • Semantic retrieval
  • Gemini integration
  • Structured technical responses
  • Source/chunk inspection
  • Retrieval telemetry
  • Streamlit interface
  • Docker containerization
  • Local Docker execution

Future Improvements

  • Improve container health-check configuration
  • Add automated testing
  • Add CI/CD
  • Improve retrieval evaluation
  • Add hybrid search
  • Add reranking
  • Add conversation memory
  • Deploy to a cloud platform

πŸ“œ License

This project is licensed under the MIT License.

About

A lightweight RAG assistant built to prepare for Cloud and DevOps interviews. It runs a local FAISS index for sub-second retrieval across Docker, Kubernetes, AWS, and Linux notes, paired with Gemini Flash for grounded, interview-ready answers. Fully containerized with Docker.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages