A production-grade, containerized Retrieval-Augmented Generation (RAG) system engineered to deliver low-latency, interview-grade explanations across core Cloud, DevOps, and Infrastructure domains.
Cloud//Base combines a curated Markdown knowledge base, local HuggingFace embeddings, FAISS semantic retrieval, Gemini, and a Streamlit interface to provide grounded technical answers with transparent source inspection and retrieval telemetry.
The main Cloud//Base interface provides system status, supported knowledge domains, retrieval depth, suggested interview topics, and the primary question interface.
Cloud//Base transforms retrieved knowledge into a structured technical explanation designed for both learning and interview preparation.
Each query exposes runtime telemetry including total latency, local vector retrieval latency, Gemini generation time, and the number of retrieved context chunks.
The retrieved knowledge chunks can be inspected directly, including their originating Markdown file, similarity score, and retrieved content.
The application is packaged and executed as a Docker container, providing a reproducible runtime environment for Cloud//Base.
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Knowledge Base (.md) β
β AWS β’ Docker β’ Kubernetes β’ Linux β’ Terraform β’ Networking β
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β
β MarkdownHeaderTextSplitter
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Local HuggingFace Embeddings β
β all-MiniLM-L6-v2 β’ CPU β’ βΉ0 Cost β
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β
β Normalized Vectors
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β FAISS Vector Store β
β Local Cosine Distance Similarity Index β
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β
β Top-K Semantic Chunks
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Gemini 3.5 Flash Lite LLM β
β Grounded Prompt Synthesis + Fallback β
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Streamlit Web Interface β
β Dark UI + Real-Time Telemetry β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Cloud//Base follows a retrieval-first architecture:
User Question
β
βΌ
Question Embedding
β
βΌ
FAISS Similarity Search
β
βΌ
Top-K Relevant Chunks
β
βΌ
Retrieved Context + User Question
β
βΌ
Grounded Gemini Prompt
β
βΌ
Structured Technical Answer
β
ββββββββββββββββΊ Retrieved Sources
β
ββββββββββββββββΊ Retrieval Telemetry
The system separates the retrieval stage from the generation stage, allowing the retrieved context and source information to be inspected alongside the generated answer.
Cloud//Base retrieves relevant information from a curated Cloud, DevOps, and Infrastructure knowledge base before generating an answer.
This allows responses to be grounded in the project's own technical documentation.
FAISS is used as the local vector store for semantic similarity search.
The local retrieval pipeline avoids the need for an externally hosted vector database.
Embeddings are generated locally using:
sentence-transformers/all-MiniLM-L6-v2
The embedding model runs locally on CPU without requiring a paid embedding API.
The knowledge base is organized into Markdown files covering:
- AWS
- Docker
- Kubernetes
- Linux
- Networking
- Terraform
Responses are organized around technical sections such as:
- Executive Definition & Core Value
- Architecture & Technical Mechanics
- Production Best Practices
- Interview Q&A / Key Takeaway
This makes the application useful for both technical learning and interview preparation.
Retrieved chunks can be inspected directly.
The interface exposes information including:
- Source Markdown file
- Similarity score
- Retrieved content
- Number of retrieved chunks
Each query provides runtime information including:
- Total latency
- Local retrieval latency
- Gemini generation latency
- Number of context chunks retrieved
The application can be packaged into a Docker image and run as a container, providing a consistent runtime environment independent of the local Python installation.
| Layer | Technology |
|---|---|
| Language | Python |
| UI | Streamlit |
| RAG Framework | LangChain |
| LLM | Google Gemini |
| Embeddings | HuggingFace Sentence Transformers |
| Embedding Model | all-MiniLM-L6-v2 |
| Vector Store | FAISS |
| Knowledge Format | Markdown |
| Containerization | Docker |
| Version Control | Git / GitHub |
cloud-knowledge-assistant/
β
βββ app/
β βββ __init__.py
β βββ embeddings.py # HuggingFace & FAISS index management
β βββ ingestion.py # Markdown header chunking pipeline
β βββ main.py # Streamlit UI & telemetry dashboard
β βββ rag.py # LangChain prompt chaining & Gemini routing
β βββ retrieval.py # Similarity search helper utilities
β
βββ assets/
β βββ 01-landing-ui.png
β βββ 02-query-response.png
β βββ 03-telemetry-metrics.png
β βββ 04-grounded-sources.png
β βββ 05-docker-running.png
β
βββ data/
β βββ vectorstore/
β βββ ...
β
βββ knowledge/
β βββ aws.md
β βββ docker.md
β βββ kubernetes.md
β βββ linux.md
β βββ networking.md
β βββ terraform.md
β
βββ .dockerignore
βββ .env.example
βββ .gitignore
βββ Dockerfile
βββ requirements.txt
βββ README.md
git clone https://github.com/<your-username>/cloud-knowledge-assistant.git
cd cloud-knowledge-assistantCreate a .env file in the project root:
GEMINI_API_KEY="your_google_ai_studio_api_key"
GEMINI_MODEL="gemini-3.5-flash-lite"
EMBEDDING_MODEL_NAME="sentence-transformers/all-MiniLM-L6-v2"
VECTOR_STORE_DIR="data/vectorstore"Important: Never commit
.envor API keys to GitHub.
Use .env.example to document required environment variables without exposing credentials.
py -3.12 -m venv venv.\venv\Scripts\Activate.ps1python --versionExpected:
Python 3.12.x
pip install -r requirements.txtstreamlit run app/main.pyThe application will be available at:
http://localhost:8501
Cloud//Base can also be executed inside Docker.
docker build -t cloud-knowledge-assistant:latest .docker run -d `
-p 8501:8501 `
--env-file .env `
--name cloud-assistant `
cloud-knowledge-assistant:latestAccess the application at:
http://localhost:8501
docker psdocker logs cloud-assistantdocker stop cloud-assistantdocker rm cloud-assistantCloud//Base uses environment variables for API credentials and configuration.
Example:
.env
β
βββ GEMINI_API_KEY
βββ GEMINI_MODEL
βββ EMBEDDING_MODEL_NAME
βββ VECTOR_STORE_DIR
The real .env file should never be committed to source control.
Recommended .gitignore entries:
.env
venv/
__pycache__/
*.pycCloud//Base can be used to explore technical concepts such as:
Explain ECS vs EKS with an interview answer.
How does chmod 755 work mathematically?
What is the difference between Layer 4 and Layer 7 Load Balancers?
How does Terraform state locking work with DynamoDB?
What is Docker?
What is a Kubernetes Pod?
What is the role of etcd in Kubernetes?
What is a Docker image vs a Docker container?
The retrieved context is shown alongside the generated response so that the grounding of the answer can be inspected.
For each question, Cloud//Base tracks:
| Metric | Description |
|---|---|
| Total Latency | Total time taken to process the request |
| Local Retrieval | Time spent performing FAISS retrieval |
| Gemini Generation | Time spent generating the final response |
| Context Chunks | Number of knowledge chunks retrieved |
Example telemetry from a query:
TOTAL LATENCY 4.965s
LOCAL RETRIEVAL 0.3274s
GEMINI GENERATION 4.638s
CONTEXT CHUNKS 4 retrieved
The objective of Cloud//Base is to build a practical RAG application where the underlying retrieval and generation pipeline can be understood and inspected rather than treating the LLM as a black box.
The project demonstrates practical experience with:
Python
β
βββ LangChain
β
βββ RAG
β
βββ Embeddings
β
βββ Vector Search
β
βββ FAISS
β
βββ Gemini
β
βββ Streamlit
β
βΌ
Docker
- Curated Cloud/DevOps knowledge base
- Markdown document ingestion
- Markdown header-based chunking
- Local HuggingFace embeddings
- FAISS vector store
- Semantic retrieval
- Gemini integration
- Structured technical responses
- Source/chunk inspection
- Retrieval telemetry
- Streamlit interface
- Docker containerization
- Local Docker execution
- Improve container health-check configuration
- Add automated testing
- Add CI/CD
- Improve retrieval evaluation
- Add hybrid search
- Add reranking
- Add conversation memory
- Deploy to a cloud platform
This project is licensed under the MIT License.




