Skip to content

Latest commit

ย 

History

21 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐ŸŽ“ Student Score Prediction โ€” End-to-End Machine Learning Project

Python Version Framework ML Libraries License Status

An end-to-end production-grade Machine Learning application to predict student Maths Scores based on demographic factors, educational backgrounds, and other academic performance indicators. Built with an automated data pipeline, robust preprocessing, hyperparameter optimization, custom logging/exception handling, and an interactive Flask web application.


๐Ÿ“Œ Table of Contents


๐Ÿ“– Overview

In modern education, understanding the factors that influence academic performance is crucial for educators, parents, and policymakers. This project implements a full lifecycle Machine Learning solution:

  • Comprehensive Exploratory Data Analysis (EDA) and visualization.
  • Automated ETL pipeline for data ingestion, validation, and train-test splitting.
  • Multi-step preprocessing (missing value imputation, standard scaling, and one-hot encoding).
  • Model evaluation across multiple regression algorithms with grid-search hyperparameter tuning.
  • Interactive user interface built with Flask for real-time score prediction.
  • Enterprise-grade modular code architecture adhering to best software engineering practices.

๐ŸŽฏ Problem Statement

The objective is to predict a student's Math Score (0 - 100) by analyzing:

  1. Demographic attributes (Gender, Ethnicity).
  2. Socioeconomic factors (Parental Level of Education, Lunch Type).
  3. Academic preparation (Test Preparation Course completion).
  4. Correlated academic scores (Reading Score, Writing Score).

๐Ÿ“Š Dataset & Features

The dataset comprises student examination records with demographic and test result variables:

Feature Name Data Type Type Description / Values
gender Categorical Input male, female
race_ethnicity Categorical Input group A, group B, group C, group D, group E
parental_level_of_education Categorical Input some high school, high school, some college, associate's degree, bachelor's degree, master's degree
lunch Categorical Input standard, free/reduced
test_preparation_course Categorical Input none, completed
reading_score Numerical Input Student's reading test score ($0 - 100$)
writing_score Numerical Input Student's writing test score ($0 - 100$)
math_score Numerical Target Student's math test score ($0 - 100$)

๐Ÿ—๏ธ Project Architecture & Workflow

The pipeline executes through a sequence of decoupled modules:

flowchart TD
    A[Raw Dataset stud.xlsx] --> B[Data Ingestion Component]
    B -->|Splits 80/20| C[artifacts/train.csv & test.csv]
    B --> D[artifacts/raw.csv]
    
    C --> E[Data Transformation Component]
    E -->|SimpleImputer + StandardScaler| F[Numerical Pipeline]
    E -->|SimpleImputer + OneHotEncoder| G[Categorical Pipeline]
    F & G --> H[ColumnTransformer Object]
    H --> I[artifacts/preprocessor.pkl]
    H --> J[Transformed Arrays train_arr & test_arr]
    
    J --> K[Model Trainer Component]
    K -->|GridSearchCV Tuning| L[Model Evaluation Benchmarking]
    L --> M[Select Best Regressor Model]
    M --> N[artifacts/model.pkl]
    
    O[User Input via Flask Web UI] --> P[CustomData / Predict Pipeline]
    I --> P
    N --> P
    P --> Q[Predicted Math Score Result]
Loading

๐Ÿ“‚ Directory Structure

mlproject/
โ”œโ”€โ”€ .gitignore                      # Git ignore file
โ”œโ”€โ”€ app.py                          # Flask Web Application & API routes
โ”œโ”€โ”€ main.py                         # Application entrypoint
โ”œโ”€โ”€ requirements.txt                # Project dependencies
โ”œโ”€โ”€ setup.py                        # Package distribution script
โ”œโ”€โ”€ pyproject.toml                  # Build tool configuration
โ”œโ”€โ”€ uv.lock                         # Lockfile for dependency management
โ”‚
โ”œโ”€โ”€ artifacts/                      # Generated pipeline artifacts (persisted objects)
โ”‚   โ”œโ”€โ”€ raw.csv                     # Ingested raw dataset
โ”‚   โ”œโ”€โ”€ train.csv                   # Training split (80%)
โ”‚   โ”œโ”€โ”€ test.csv                    # Testing split (20%)
โ”‚   โ”œโ”€โ”€ preprocessor.pkl            # Serialized preprocessor pipeline
โ”‚   โ””โ”€โ”€ model.pkl                   # Serialized best trained ML model
โ”‚
โ”œโ”€โ”€ logs/                           # Auto-generated execution logs
โ”‚   โ””โ”€โ”€ <timestamp>.log             # Formatted logs with tracebacks
โ”‚
โ”œโ”€โ”€ notebook/                       # Jupyter Notebooks for exploration
โ”‚   โ”œโ”€โ”€ data/
โ”‚   โ”‚   โ””โ”€โ”€ stud.xlsx               # Source raw dataset
โ”‚   โ”œโ”€โ”€ 1 . EDA STUDENT PERFORMANCE.ipynb # Exploratory Data Analysis & Visualizations
โ”‚   โ””โ”€โ”€ 2. MODEL TRAINING.ipynb     # Model experimentation & prototyping
โ”‚
โ”œโ”€โ”€ src/                            # Modular source code
โ”‚   โ”œโ”€โ”€ __init__.py                 # Package marker
โ”‚   โ”œโ”€โ”€ exception.py                # Custom exception handler with detailed trace
โ”‚   โ”œโ”€โ”€ logger.py                   # Centralized logging configuration
โ”‚   โ”œโ”€โ”€ utils.py                    # Utility helpers (save_object, evaluate_models, etc.)
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ components/                 # Core pipeline components
โ”‚   โ”‚   โ”œโ”€โ”€ __init__.py
โ”‚   โ”‚   โ”œโ”€โ”€ data_ingestions.py      # Data ingestion & train/test partitioning
โ”‚   โ”‚   โ”œโ”€โ”€ data_transformation.py  # Preprocessing pipelines & feature encoding
โ”‚   โ”‚   โ””โ”€โ”€ model_trainer.py        # Model benchmarking & hyperparameter optimization
โ”‚   โ”‚
โ”‚   โ””โ”€โ”€ pipeline/                   # Execution pipelines
โ”‚       โ”œโ”€โ”€ __init__.py
โ”‚       โ”œโ”€โ”€ train_pipeline.py       # End-to-end training pipeline orchestrator
โ”‚       โ””โ”€โ”€ predict_pipeline.py     # Inference pipeline with custom data wrapper
โ”‚
โ””โ”€โ”€ templates/                      # Flask HTML templates
    โ”œโ”€โ”€ index.html                  # Welcome landing page
    โ””โ”€โ”€ home.html                   # Prediction input form & results display

โš™๏ธ Core Components

1. Data Ingestion

  • File: src/components/data_ingestions.py
  • Loads raw data from notebook/data/stud.xlsx.
  • Validates directories and saves raw.csv under artifacts/.
  • Performs an 80/20 stratified train-test split (random_state=42) and outputs train.csv and test.csv.

2. Data Transformation

  • File: src/components/data_transformation.py
  • Constructs specialized preprocessing pipelines:
    • Numerical Features (reading_score, writing_score): Missing value imputation via median, followed by StandardScaler.
    • Categorical Features (gender, race_ethnicity, parental_level_of_education, lunch, test_preparation_course): Missing value imputation via most_frequent, followed by OneHotEncoder.
  • Bundles features using ColumnTransformer.
  • Fits on training data, transforms both train and test splits, and serializes the transformer to artifacts/preprocessor.pkl.

3. Model Training & Hyperparameter Tuning

  • File: src/components/model_trainer.py
  • Evaluates multiple regression algorithms with GridSearchCV (3-fold cross-validation, R^2 scoring):
    • Decision Tree Regressor (criterion, max_depth, min_samples_split, min_samples_leaf)
    • Random Forest Regressor (n_estimators, max_depth, min_samples_split, min_samples_leaf)
    • Gradient Boosting Regressor (n_estimators, learning_rate, max_depth, subsample)
    • Linear Regression
    • XGBoost Regressor (n_estimators, learning_rate, max_depth, subsample)
    • CatBoost Regressor (iterations, depth, learning_rate)
    • AdaBoost Regressor (n_estimators, learning_rate)
  • Selects the top-performing model exceeding the score threshold (R^2 \ge 0.6) and serializes it to artifacts/model.pkl.

4. Prediction Pipeline

  • File: src/pipeline/predict_pipeline.py
  • Exposes:
    • CustomData: Maps input features from web requests or APIs into structured pandas.DataFrame.
    • PredicPipeline: Loads artifacts/preprocessor.pkl and artifacts/model.pkl to generate predictions for new data points.

5. Exception Handling & Logging

  • Files: src/exception.py, src/logger.py
  • Custom CustomException formats detailed error strings with exact file names and line numbers.
  • Centralized logger writes runtime events with timestamps and log levels into the logs/ directory.

๐Ÿค– Machine Learning Models & Evaluation

The training framework systematically compares multiple regressors using $R^2$ Score (Coefficient of Determination):

       Rยฒ Score = 1 - (SS_res / SS_tot)
Model Hyperparameter Search Space
Linear Regression Default baseline
Decision Tree criterion, max_depth: [None, 5, 10], min_samples_split: [2, 5]
Random Forest n_estimators: [50, 100], max_depth: [None, 10], min_samples_split: [2, 5]
Gradient Boosting n_estimators: [50, 100], learning_rate: [0.05, 0.1], subsample: [0.8, 1.0]
XGBoost n_estimators: [50, 100], learning_rate: [0.05, 0.1], max_depth: [3, 5]
CatBoost iterations: [100, 200], depth: [4, 6], learning_rate: [0.05, 0.1]
AdaBoost n_estimators: [50, 100], learning_rate: [0.05, 0.1]

๐Ÿš€ Installation & Setup

Prerequisites

  • Python 3.8 to 3.12 installed on your system.
  • Git installed.

1. Clone the Repository

git clone https://github.com/alphacrypto246/Student-Score-Prediction.git
cd Student-Score-Prediction

2. Create and Activate a Virtual Environment

On Windows (PowerShell):

python -m venv .venv
.venv\Scripts\Activate.ps1

On Linux / macOS:

python3 -m venv .venv
source .venv/bin/activate

3. Install Dependencies

pip install --upgrade pip
pip install -r requirements.txt

(Optional) Install the package in editable mode:

pip install -e .

๐Ÿ’ป Usage Guide

Running the Training Pipeline

To execute the end-to-end data ingestion, transformation, and model training workflow:

python src/components/data_ingestions.py

This will:

  1. Ingest notebook/data/stud.xlsx.
  2. Generate train.csv, test.csv, and raw.csv in artifacts/.
  3. Fit and save the preprocessing pipeline to artifacts/preprocessor.pkl.
  4. Train and tune all regression models, output the best $R^2$ score, and save the best model to artifacts/model.pkl.

Running the Web Application

Launch the Flask development server:

python app.py

Once running, access the web application in your browser at:

  • Home / Welcome: http://127.0.0.1:5000/
  • Prediction Form: http://127.0.0.1:5000/predictdata

๐ŸŒ Web Interface Demo

  1. Navigate to http://127.0.0.1:5000/predictdata.
  2. Enter student attributes:
    • Select Gender, Ethnicity, and Parental Level of Education.
    • Select Lunch Type and Test Preparation Course.
    • Input Reading Score ($0-100$) and Writing Score ($0-100$).
  3. Click Predict your Maths Score.
  4. The estimated Math Score is calculated and displayed instantly on screen.

๐Ÿ› ๏ธ Technologies & Tools

Category Technologies
Core Language Python 3.10+
Web Framework Flask, Jinja2, HTML5/CSS3
Machine Learning Scikit-Learn, XGBoost, CatBoost
Data Processing Pandas, NumPy, OpenPyXL
Data Visualization Matplotlib, Seaborn
Model Serialization Pickle, Dill
Package Management Setuptools, pip, uv

๐Ÿ”ฎ Future Enhancements

  • Add Docker containerization (Dockerfile and docker-compose.yml) for seamless deployment.
  • Build CI/CD workflow with GitHub Actions.
  • Deploy to cloud platforms (AWS Elastic Beanstalk / Azure App Service / Render / Hugging Face Spaces).
  • Add unit tests and integration tests using pytest.
  • Implement model monitoring and MLflow / DVC experiment tracking.

๐Ÿ‘ค Author & Acknowledgements


๐Ÿ“„ License

This project is licensed under the MIT License โ€” see the LICENSE file for details.

About

Production-grade end-to-end ML pipeline for student score prediction featuring modular ETL, custom logging/exception handling, model hyperparameter tuning, and Flask deployment.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages