Skip to content
View rutu6103's full-sized avatar

Block or report rutu6103

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rutu6103/README.md

Hi, I'm Rutuja Kadam

Data Scientist and Analyst with an M.Sc. in Statistics. I turn complex, multi-source data into validated metrics, semantic models, and decision-ready dashboards.

Currently open to Data Scientist, Data Analyst, Business Intelligence, and analytics-focused Data Engineering roles.

What I do

Machine Learning and Statistics - Regression, classification, clustering, hypothesis testing, time-series analysis, feature engineering, and model evaluation. Translate statistical results into business language.

Business Intelligence - Design and build Power BI reports, semantic models, and DAX measures. Import from PostgreSQL, MySQL, and file-based sources. Publish to Power BI Service with stakeholder access.

Data Engineering Foundations - Python-based ETL, API ingestion, dimensional modeling, SQL analytics, and data-quality validation. Build pipelines that are documented and reproducible.

Applied AI Evaluation - Prompt and model comparison, LLM output evaluation, taxonomy validation, and structured-tag extraction from unstructured text.

Tech stack

Category Tools
Languages Python, SQL, R
Data manipulation pandas, NumPy, scikit-learn
Visualization Matplotlib, Seaborn
Databases PostgreSQL, MySQL
BI and visualization Power BI, DAX, Power Query, semantic models
Statistics and ML Regression, classification, clustering, hypothesis testing, model evaluation
Workflow Git, GitHub, Jupyter Notebook, R Markdown, Excel

Experience

Axion Ray - Data Analyst Build Python data pipelines for cleaning, deduplication, and cross-source reconciliation. Write and validate SQL metric logic across operational datasets. Evaluate LLM outputs and compare prompt and model configurations.

Dozee - Data Analytics Intern Built and validated Power BI reports, created semantic models, imported data from PostgreSQL, MySQL, Google Sheets, and Excel, and published reports to Power BI Service for controlled stakeholder access.

Featured projects

Healthcare Analytics Pipeline - End-to-end BI and data-engineering project using public synthetic FHIR data. Python extraction and transformation, PostgreSQL dimensional modeling, SQL analytics with CTEs and window functions, data-quality checks, and a five-page Power BI dashboard with city-level drillthrough.

Python · PostgreSQL · SQL · ETL · Dimensional Modeling · Power BI · DAX

Credit Customer Segmentation - Unsupervised machine learning project comparing K-Means, Hierarchical Clustering, and DBSCAN on the South German Credit dataset, with post-hoc credit-risk interpretation.

Clustering · Unsupervised ML · Python · scikit-learn · Customer Analytics

Statistical Modeling of Product Choice - Research-oriented statistical modeling project examining factors associated with product choice. Emphasis on interpretable analysis, responsible communication, and evidence-based conclusions.

Statistical Modeling · Research · Hypothesis Testing · R · Interpretability

Currently exploring

End-to-end data science workflows including ML deployment patterns, LLM output evaluation at scale, and applied statistical modeling on real-world datasets. Learning in public through portfolio projects on this profile.

How I work

  • Business question first. I start with what decision the analysis supports, not with which tool to use.
  • Reproducible by default. Analysis is scripted, documented, and rerunnable. Results trace back to source.
  • Validation before reporting. Metrics are checked before they reach a dashboard or a stakeholder.
  • Honest about limitations. Assumptions, data-quality issues, and uncertainty are documented alongside results, not hidden underneath them.
  • Communication matters. A finding only a statistician can interpret is half-finished.

Connect

Based in Bengaluru, India. Open to on-site and hybrid roles in Bengaluru, Mumbai, and Pune.

Pinned Loading

  1. healthcare-analytics-pipeline healthcare-analytics-pipeline Public

    End-to-end healthcare BI project using synthetic FHIR data, Python ETL, PostgreSQL dimensional modelling, SQL analytics, data-quality checks, and Power BI.

    Python 1

  2. credit-customer-segmentation credit-customer-segmentation Public

    Banking-focused customer segmentation using K-Means, Hierarchical Clustering, and DBSCAN on the South German Credit dataset, with post-hoc credit-risk analysis.

    Jupyter Notebook 1

  3. menstrual-hygiene-choice-statistical-modeling menstrual-hygiene-choice-statistical-modeling Public

    Exploratory statistical modeling of menstrual-hygiene awareness, product preferences, and switching behavior using survey data in R.

    R 1