Skip to content
View sumitstat07's full-sized avatar
🌌
Statistics and Data science
🌌
Statistics and Data science

Block or report sumitstat07

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sumitstat07/README.md

Sumit Sana

M.S Statistics & Data Science (SDS) @ IIT Kanpur  |  Data Science · AI/ML · Deep Learning · Statistics · Quant · Data Analytics

Email LinkedIn GitHub followers


🧭 About Me

  • 🎓 Pursuing M.S SDS, Department of Statistics & Data Science, IIT Kanpur
  • 📊 I build statistically-grounded ML systems — not just models, but pipelines that hold up to validation: time-series diagnostics, backtests, and out-of-sample checks
  • 💼 Interested in Data Science, AI/ML, Deep Learning, Statistics, Quantitative Analysis, and Data Analytics roles
  • 🔭 Currently deepening my work in time-series econometrics and statistical machine learning
  • 🌱 Always exploring how classical statistics (hypothesis testing, stochastic processes, econometrics) and modern ML complement each other
  • 📫 Reach me at sumits25@iitk.ac.in or connect on LinkedIn

🚀 Featured Projects

Project What it does Stack
🛡️ SpectraShield Detects transaction fraud by uncovering a 24-hour spectral rhythm in the data (ADF/KPSS, Welch PSD) — XGBoost + SHAP, AUC 0.80 Python XGBoost SHAP Streamlit
🏦 Mortgage Credit Risk Modeling PD/LGD/EAD credit risk framework on Freddie Mac loan data, validated against realized 2007-crisis losses Logistic Regression XGBoost IFRS9
📈 RBI Sentiment & Bond Yields Scores hawkish/dovish tone across 163 RBI policy documents with FinBERT and links it to bond yield moves (69% directional accuracy) FinBERT Econometrics NLP
🖼️ CNN for CIFAR-10 Classification Deep CNN built from scratch in PyTorch to classify 60,000 images across 10 categories (74.71% test accuracy) PyTorch Deep Learning CNN
🍔 Big Mac Index Analysis SQL + Power BI analysis of currency valuation across 57 countries using window functions MySQL Power BI DAX
📊 Black–Litterman Portfolio Optimization Portfolio allocation framework using oil-price-shock-derived views, validated with walk-forward backtesting Python Quant Finance
🎯 Stock Market Anomaly Detection Z-score and ARIMA-based anomaly detection for NSE stocks with a live R Shiny dashboard R Shiny ARIMA
🔎 More projects — resume screening, workout coaching, biometric attendance, and deep learning — on my repositories page.

🛠️ Tech Stack

Languages Python R SQL LaTeX

ML / Data PyTorch scikit--learn XGBoost Pandas NumPy

Visualization / Apps Streamlit Power BI Shiny

Tools Git FastAPI Supabase


💬 Always happy to talk time series, credit risk modeling, or quant finance — feel free to reach out!

Pinned Loading

  1. RepLica RepLica Public

    RepLica — an AI-powered real-time workout coach that tracks exercise form via webcam (MediaPipe pose detection) and gives live voice feedback using an LLM (Groq) and TTS

    Python

  2. verita-main verita-main Public

    Core production engine for Verita AI. Implements deep-metric facial embeddings (Dlib/FaceRecognition) and sequential audio biometrics over a secure Supabase persistence layer.

    Python

  3. creditwise-loan-approval-ml creditwise-loan-approval-ml Public

    ML pipeline predicting loan approval using Logistic Regression, KNN & Naive Bayes

    Jupyter Notebook

  4. Resumate Resumate Public

    AI-powered resume ATS scorer — matches your resume against a job description using NLP and LLM feedback. Built with FastAPI, Streamlit, spaCy & Sentence Transformers.

    Jupyter Notebook

  5. RNN-for-sentiment-analysis-of-IMDB-data RNN-for-sentiment-analysis-of-IMDB-data Public

    RNN-based sentiment classifier trained on the IMDB movie reviews dataset. Includes text preprocessing (cleaning, stopword removal, stemming), TF-IDF vectorization, and a PyTorch RNN model for binar…

    Jupyter Notebook

  6. smartcart-customer-segmentation smartcart-customer-segmentation Public

    Unsupervised customer segmentation on e-commerce data using K-Means and Agglomerative Clustering

    Jupyter Notebook 2