Computer Science student focused on Data Science and Machine Learning, based in São José dos Campos, Brazil. I build end-to-end local AI: private LLM inference, RAG, computer vision and OCR, async Python pipelines that hold up under load, and the full-stack interfaces people actually use.
I take part in AI research, have optimized real e-commerce platforms, and write about what I build on Medium.
| Project | What it is | Stack |
|---|---|---|
| PolyRAG | Local federated RAG that routes across graph, vector and SQL stores by math, not by an agent | FastAPI · Qdrant · llama.cpp · SvelteKit |
| Voice Assistant | 100% local voice AI on an AMD GPU via Vulkan, no CUDA and no cloud | whisper.cpp · llama.cpp · GGML |
| RasterScope | Satellite land-cover change with uncertainty maps · demo | U-Net · ONNX · FastAPI · React |
| AeroPulse | Turbofan predictive maintenance with remaining-useful-life estimates | XGBoost · FastAPI · React · Docker |
| ToolHaven | A toolbox app for Windows · site | Tauri · Rust · React · FFmpeg |
| Manim Editor | Visual editor for math animations, no code needed · demo | Tauri · React · Python · Manim |
More case studies at diegomirhan.com/en/projects.
AI tooling · Ollama · vLLM · llama.cpp · Hugging Face Transformers · AWQ quantization · Qdrant · LangChain · CrewAI · ONNX



