Data Scientist and Analyst with an M.Sc. in Statistics. I turn complex, multi-source data into validated metrics, semantic models, and decision-ready dashboards.
Currently open to Data Scientist, Data Analyst, Business Intelligence, and analytics-focused Data Engineering roles.
Machine Learning and Statistics - Regression, classification, clustering, hypothesis testing, time-series analysis, feature engineering, and model evaluation. Translate statistical results into business language.
Business Intelligence - Design and build Power BI reports, semantic models, and DAX measures. Import from PostgreSQL, MySQL, and file-based sources. Publish to Power BI Service with stakeholder access.
Data Engineering Foundations - Python-based ETL, API ingestion, dimensional modeling, SQL analytics, and data-quality validation. Build pipelines that are documented and reproducible.
Applied AI Evaluation - Prompt and model comparison, LLM output evaluation, taxonomy validation, and structured-tag extraction from unstructured text.
| Category | Tools |
|---|---|
| Languages | Python, SQL, R |
| Data manipulation | pandas, NumPy, scikit-learn |
| Visualization | Matplotlib, Seaborn |
| Databases | PostgreSQL, MySQL |
| BI and visualization | Power BI, DAX, Power Query, semantic models |
| Statistics and ML | Regression, classification, clustering, hypothesis testing, model evaluation |
| Workflow | Git, GitHub, Jupyter Notebook, R Markdown, Excel |
Axion Ray - Data Analyst Build Python data pipelines for cleaning, deduplication, and cross-source reconciliation. Write and validate SQL metric logic across operational datasets. Evaluate LLM outputs and compare prompt and model configurations.
Dozee - Data Analytics Intern Built and validated Power BI reports, created semantic models, imported data from PostgreSQL, MySQL, Google Sheets, and Excel, and published reports to Power BI Service for controlled stakeholder access.
Healthcare Analytics Pipeline - End-to-end BI and data-engineering project using public synthetic FHIR data. Python extraction and transformation, PostgreSQL dimensional modeling, SQL analytics with CTEs and window functions, data-quality checks, and a five-page Power BI dashboard with city-level drillthrough.
Python · PostgreSQL · SQL · ETL · Dimensional Modeling · Power BI · DAX
Credit Customer Segmentation - Unsupervised machine learning project comparing K-Means, Hierarchical Clustering, and DBSCAN on the South German Credit dataset, with post-hoc credit-risk interpretation.
Clustering · Unsupervised ML · Python · scikit-learn · Customer Analytics
Statistical Modeling of Product Choice - Research-oriented statistical modeling project examining factors associated with product choice. Emphasis on interpretable analysis, responsible communication, and evidence-based conclusions.
Statistical Modeling · Research · Hypothesis Testing · R · Interpretability
End-to-end data science workflows including ML deployment patterns, LLM output evaluation at scale, and applied statistical modeling on real-world datasets. Learning in public through portfolio projects on this profile.
- Business question first. I start with what decision the analysis supports, not with which tool to use.
- Reproducible by default. Analysis is scripted, documented, and rerunnable. Results trace back to source.
- Validation before reporting. Metrics are checked before they reach a dashboard or a stakeholder.
- Honest about limitations. Assumptions, data-quality issues, and uncertainty are documented alongside results, not hidden underneath them.
- Communication matters. A finding only a statistician can interpret is half-finished.
Based in Bengaluru, India. Open to on-site and hybrid roles in Bengaluru, Mumbai, and Pune.