An end-to-end production-grade Machine Learning application to predict student Maths Scores based on demographic factors, educational backgrounds, and other academic performance indicators. Built with an automated data pipeline, robust preprocessing, hyperparameter optimization, custom logging/exception handling, and an interactive Flask web application.
- Overview
- Problem Statement
- Dataset & Features
- Project Architecture & Workflow
- Directory Structure
- Core Components
- Machine Learning Models & Evaluation
- Installation & Setup
- Usage Guide
- Web Interface Demo
- Technologies & Tools
- Future Enhancements
- Author & Acknowledgements
In modern education, understanding the factors that influence academic performance is crucial for educators, parents, and policymakers. This project implements a full lifecycle Machine Learning solution:
- Comprehensive Exploratory Data Analysis (EDA) and visualization.
- Automated ETL pipeline for data ingestion, validation, and train-test splitting.
- Multi-step preprocessing (missing value imputation, standard scaling, and one-hot encoding).
- Model evaluation across multiple regression algorithms with grid-search hyperparameter tuning.
- Interactive user interface built with Flask for real-time score prediction.
- Enterprise-grade modular code architecture adhering to best software engineering practices.
The objective is to predict a student's Math Score (0 - 100) by analyzing:
- Demographic attributes (Gender, Ethnicity).
- Socioeconomic factors (Parental Level of Education, Lunch Type).
- Academic preparation (Test Preparation Course completion).
- Correlated academic scores (Reading Score, Writing Score).
The dataset comprises student examination records with demographic and test result variables:
| Feature Name | Data Type | Type | Description / Values |
|---|---|---|---|
gender |
Categorical | Input |
male, female
|
race_ethnicity |
Categorical | Input |
group A, group B, group C, group D, group E
|
parental_level_of_education |
Categorical | Input |
some high school, high school, some college, associate's degree, bachelor's degree, master's degree
|
lunch |
Categorical | Input |
standard, free/reduced
|
test_preparation_course |
Categorical | Input |
none, completed
|
reading_score |
Numerical | Input | Student's reading test score ( |
writing_score |
Numerical | Input | Student's writing test score ( |
math_score |
Numerical | Target | Student's math test score ( |
The pipeline executes through a sequence of decoupled modules:
flowchart TD
A[Raw Dataset stud.xlsx] --> B[Data Ingestion Component]
B -->|Splits 80/20| C[artifacts/train.csv & test.csv]
B --> D[artifacts/raw.csv]
C --> E[Data Transformation Component]
E -->|SimpleImputer + StandardScaler| F[Numerical Pipeline]
E -->|SimpleImputer + OneHotEncoder| G[Categorical Pipeline]
F & G --> H[ColumnTransformer Object]
H --> I[artifacts/preprocessor.pkl]
H --> J[Transformed Arrays train_arr & test_arr]
J --> K[Model Trainer Component]
K -->|GridSearchCV Tuning| L[Model Evaluation Benchmarking]
L --> M[Select Best Regressor Model]
M --> N[artifacts/model.pkl]
O[User Input via Flask Web UI] --> P[CustomData / Predict Pipeline]
I --> P
N --> P
P --> Q[Predicted Math Score Result]
mlproject/
โโโ .gitignore # Git ignore file
โโโ app.py # Flask Web Application & API routes
โโโ main.py # Application entrypoint
โโโ requirements.txt # Project dependencies
โโโ setup.py # Package distribution script
โโโ pyproject.toml # Build tool configuration
โโโ uv.lock # Lockfile for dependency management
โ
โโโ artifacts/ # Generated pipeline artifacts (persisted objects)
โ โโโ raw.csv # Ingested raw dataset
โ โโโ train.csv # Training split (80%)
โ โโโ test.csv # Testing split (20%)
โ โโโ preprocessor.pkl # Serialized preprocessor pipeline
โ โโโ model.pkl # Serialized best trained ML model
โ
โโโ logs/ # Auto-generated execution logs
โ โโโ <timestamp>.log # Formatted logs with tracebacks
โ
โโโ notebook/ # Jupyter Notebooks for exploration
โ โโโ data/
โ โ โโโ stud.xlsx # Source raw dataset
โ โโโ 1 . EDA STUDENT PERFORMANCE.ipynb # Exploratory Data Analysis & Visualizations
โ โโโ 2. MODEL TRAINING.ipynb # Model experimentation & prototyping
โ
โโโ src/ # Modular source code
โ โโโ __init__.py # Package marker
โ โโโ exception.py # Custom exception handler with detailed trace
โ โโโ logger.py # Centralized logging configuration
โ โโโ utils.py # Utility helpers (save_object, evaluate_models, etc.)
โ โ
โ โโโ components/ # Core pipeline components
โ โ โโโ __init__.py
โ โ โโโ data_ingestions.py # Data ingestion & train/test partitioning
โ โ โโโ data_transformation.py # Preprocessing pipelines & feature encoding
โ โ โโโ model_trainer.py # Model benchmarking & hyperparameter optimization
โ โ
โ โโโ pipeline/ # Execution pipelines
โ โโโ __init__.py
โ โโโ train_pipeline.py # End-to-end training pipeline orchestrator
โ โโโ predict_pipeline.py # Inference pipeline with custom data wrapper
โ
โโโ templates/ # Flask HTML templates
โโโ index.html # Welcome landing page
โโโ home.html # Prediction input form & results display
- File:
src/components/data_ingestions.py - Loads raw data from
notebook/data/stud.xlsx. - Validates directories and saves
raw.csvunderartifacts/. - Performs an 80/20 stratified train-test split (
random_state=42) and outputstrain.csvandtest.csv.
- File:
src/components/data_transformation.py - Constructs specialized preprocessing pipelines:
- Numerical Features (
reading_score,writing_score): Missing value imputation viamedian, followed byStandardScaler. - Categorical Features (
gender,race_ethnicity,parental_level_of_education,lunch,test_preparation_course): Missing value imputation viamost_frequent, followed byOneHotEncoder.
- Numerical Features (
- Bundles features using
ColumnTransformer. - Fits on training data, transforms both train and test splits, and serializes the transformer to
artifacts/preprocessor.pkl.
- File:
src/components/model_trainer.py - Evaluates multiple regression algorithms with
GridSearchCV(3-fold cross-validation, R^2 scoring):- Decision Tree Regressor (criterion, max_depth, min_samples_split, min_samples_leaf)
- Random Forest Regressor (n_estimators, max_depth, min_samples_split, min_samples_leaf)
- Gradient Boosting Regressor (n_estimators, learning_rate, max_depth, subsample)
- Linear Regression
- XGBoost Regressor (n_estimators, learning_rate, max_depth, subsample)
- CatBoost Regressor (iterations, depth, learning_rate)
- AdaBoost Regressor (n_estimators, learning_rate)
- Selects the top-performing model exceeding the score threshold (R^2 \ge 0.6) and serializes it to
artifacts/model.pkl.
- File:
src/pipeline/predict_pipeline.py - Exposes:
CustomData: Maps input features from web requests or APIs into structuredpandas.DataFrame.PredicPipeline: Loadsartifacts/preprocessor.pklandartifacts/model.pklto generate predictions for new data points.
- Files:
src/exception.py,src/logger.py - Custom
CustomExceptionformats detailed error strings with exact file names and line numbers. - Centralized logger writes runtime events with timestamps and log levels into the
logs/directory.
The training framework systematically compares multiple regressors using
Rยฒ Score = 1 - (SS_res / SS_tot)
| Model | Hyperparameter Search Space |
|---|---|
| Linear Regression | Default baseline |
| Decision Tree | criterion, max_depth: [None, 5, 10], min_samples_split: [2, 5] |
| Random Forest | n_estimators: [50, 100], max_depth: [None, 10], min_samples_split: [2, 5] |
| Gradient Boosting | n_estimators: [50, 100], learning_rate: [0.05, 0.1], subsample: [0.8, 1.0] |
| XGBoost | n_estimators: [50, 100], learning_rate: [0.05, 0.1], max_depth: [3, 5] |
| CatBoost | iterations: [100, 200], depth: [4, 6], learning_rate: [0.05, 0.1] |
| AdaBoost | n_estimators: [50, 100], learning_rate: [0.05, 0.1] |
- Python 3.8 to 3.12 installed on your system.
- Git installed.
git clone https://github.com/alphacrypto246/Student-Score-Prediction.git
cd Student-Score-PredictionOn Windows (PowerShell):
python -m venv .venv
.venv\Scripts\Activate.ps1On Linux / macOS:
python3 -m venv .venv
source .venv/bin/activatepip install --upgrade pip
pip install -r requirements.txt(Optional) Install the package in editable mode:
pip install -e .To execute the end-to-end data ingestion, transformation, and model training workflow:
python src/components/data_ingestions.pyThis will:
- Ingest
notebook/data/stud.xlsx. - Generate
train.csv,test.csv, andraw.csvinartifacts/. - Fit and save the preprocessing pipeline to
artifacts/preprocessor.pkl. - Train and tune all regression models, output the best
$R^2$ score, and save the best model toartifacts/model.pkl.
Launch the Flask development server:
python app.pyOnce running, access the web application in your browser at:
- Home / Welcome:
http://127.0.0.1:5000/ - Prediction Form:
http://127.0.0.1:5000/predictdata
- Navigate to
http://127.0.0.1:5000/predictdata. - Enter student attributes:
- Select Gender, Ethnicity, and Parental Level of Education.
- Select Lunch Type and Test Preparation Course.
- Input Reading Score (
$0-100$ ) and Writing Score ($0-100$ ).
- Click Predict your Maths Score.
- The estimated Math Score is calculated and displayed instantly on screen.
| Category | Technologies |
|---|---|
| Core Language | Python 3.10+ |
| Web Framework | Flask, Jinja2, HTML5/CSS3 |
| Machine Learning | Scikit-Learn, XGBoost, CatBoost |
| Data Processing | Pandas, NumPy, OpenPyXL |
| Data Visualization | Matplotlib, Seaborn |
| Model Serialization | Pickle, Dill |
| Package Management | Setuptools, pip, uv |
- Add Docker containerization (
Dockerfileanddocker-compose.yml) for seamless deployment. - Build CI/CD workflow with GitHub Actions.
- Deploy to cloud platforms (AWS Elastic Beanstalk / Azure App Service / Render / Hugging Face Spaces).
- Add unit tests and integration tests using
pytest. - Implement model monitoring and MLflow / DVC experiment tracking.
- Author: Arya Deep Chowdhury
- Email: arya.d.chowdhury@gmail.com
- GitHub: @alphacrypto246
This project is licensed under the MIT License โ see the LICENSE file for details.