Skip to content
sisinflabPublic

Latest commit

Β 

History

2,183 Commits

Folders and files

Repository files navigation

πŸš€ WarpRec

GitHub release (latest by date) PyPI version License: MIT Python 3.12 Documentation Status PyTorch Ruff CodeCarbon MCP Powered GitHub Stars

Read the Docs

WarpRec is a flexible and efficient framework designed for building, training, evaluating, and estimating recommendation workloads. It supports a wide range of configurations, customizable pipelines, and powerful optimization tools to enhance model performance and usability.

WarpRec is designed for both beginners and experienced practitioners. For newcomers, it offers a simple and intuitive interface to explore and experiment with state-of-the-art recommendation models. For advanced users, WarpRec provides a modular and extensible architecture that allows rapid prototyping, complex experiment design, and fine-grained control over every step of the recommendation pipeline.

Whether you're learning how recommender systems work or conducting high-performance research and development, WarpRec offers the right tools to match your workflow.

πŸ—οΈ Architecture

WarpRec Architecture

WarpRec is built on 4 foundational pillars β€” Scalability, Green AI, Agentic Readiness, and Scientific Rigor β€” and organized into 5 modular engines that manage the end-to-end recommendation lifecycle:

  1. Reader β€” Ingests user-item interactions and metadata from local or cloud storage via a backend-agnostic Narwhals abstraction layer.
  2. Data Engine β€” Applies configurable filtering and splitting strategies, including cold-start protocols, to produce clean, leak-free train/validation/test sets.
  3. Recommendation Engine β€” Trains and optimizes models using PyTorch, with seamless scaling from single-GPU to multi-node Ray clusters.
  4. Evaluation Engine β€” Computes 45 GPU-accelerated metrics in a single pass with automated statistical significance testing, optional re-ranking, and inverse-propensity estimators that correct for exposure bias.
  5. Writer β€” Serializes results, checkpoints, and carbon reports to local or cloud storage.

An Application Layer serves trained models on Ray Serve, as a batched REST API and, optionally, as MCP tools for agentic AI workflows.

πŸ“š Table of Contents

✨ Key Features

  • 88 Built-in Algorithms: WarpRec ships with 88 state-of-the-art recommendation models spanning 8 paradigms β€” Unpersonalized, Content-Based, Collaborative Filtering (e.g., LightGCN, EASE$^R$, MultiVAE), Context-Aware (e.g., DeepFM, xDeepFM), Sequential (e.g., SASRec, BERT4Rec, GRU4Rec), Knowledge-Aware (e.g., KGAT, KGIN, KaHFM, KGFlex), Multimodal (e.g., VBPR, FREEDOM, BM3), and Hybrid. All models are fully configurable and extend a standardized base class, making it easy to prototype custom architectures within the same pipeline.
  • Backend-Agnostic Data Engine: Built on Narwhals, WarpRec operates over Pandas, Polars, and Spark without code changes β€” enabling a true "write-once, run-anywhere" workflow from laptop to distributed cluster. Data ingestion supports both local filesystems and cloud object storage (Azure Blob Storage).
  • Comprehensive Data Processing: The data module provides 13 filtering strategies (filter-by-rating, k-core, cold-start heuristics) and 8 splitting protocols (random/temporal Hold-Out, Leave-k-Out, Fixed Timestamp, k-fold Cross-Validation, item and user cold-start), for a total of 21 configurable strategies to ensure rigorous and reproducible experimental setups.
  • 45 GPU-Accelerated Metrics: The evaluation suite covers 45 metrics across 8 families β€” Accuracy, Rating, Coverage, Novelty, Diversity, Bias, Fairness, and Debiased β€” including multi-objective metrics for simultaneous optimization of competing goals, and inverse-propensity estimators that correct for exposure bias. All metrics are computed with full GPU acceleration for large-scale experiments.
  • Statistical Rigor: WarpRec automates hypothesis testing with paired (Student's t-test, Wilcoxon signed-rank) and independent-group (Mann-Whitney U) tests, and applies multiple comparison corrections via Bonferroni and FDR (Benjamini-Hochberg) to prevent p-hacking and ensure statistically robust conclusions.
  • Distributed Training & HPO: Seamless vertical and horizontal scaling from single-GPU to multi-node Ray clusters. Hyperparameter optimization supports Grid, Random, Bayesian, HyperOpt, Optuna, and BoHB strategies, with ASHA pruning and model-level early stopping to maximize computational efficiency.
  • Pausable & Resumable Runs: Long experiments can be stopped with a signal (Ctrl+C or SIGTERM) and resumed later from the same command. Unfinished Ray Tune trials continue from their last checkpoint and models that already completed are skipped, so a preemption on a spot instance or a cluster reclaim costs minutes rather than the whole experiment.
  • Green AI & Carbon Tracking: WarpRec is the first recommendation framework with native CodeCarbon integration, automatically quantifying energy consumption and COβ‚‚ emissions for every experiment and persisting carbon footprint reports alongside standard results.
  • Agentic AI via MCP: Served models can also be exposed as Model Context Protocol tools on the same server (server.mcp: true), so LLMs and autonomous agents call a trained recommender as a tool β€” transforming the framework from a static predictor into an interactive, agent-ready component.
  • Model Serving on Ray Serve: A model saved by the training pipeline is served with python -m warprec.serve -c serve.yml β€” a batched REST API with replicas, autoscaling and GPU placement set in configuration, and exportable to a Ray cluster or KubeRay. General, sequential, graph and context-aware models are all served from the checkpoint alone.
  • Experiment Tracking: Native integrations with Weights & Biases and MLflow for real-time monitoring of metrics, training dynamics and multi-run management, and CodeCarbon for the energy and emissions of every trial.
  • Custom Pipelines & Callbacks: Alongside the standard Training, Design, Evaluation, Swarm, and Estimate workflows, WarpRec exposes an event-driven Callback system for injecting custom logic at any stage β€” enabling complex experiments without modifying framework internals.

βš™οΈ Installation

WarpRec is designed to be easily installed via pip or via Conda. This ensures that all dependencies and the Python environment are managed consistently.

πŸš€ Quick Install (PyPI)

The easiest way to get started is using pip:

pip install warprec

WarpRec provides extra dependencies for specific use cases:

extra usage
dashboard Dashboard functionalities like MLflow and Weights & Biases.
remote-io Remote communication with cloud services like Azure.
serving Ray Serve, to serve trained models over HTTP with warprec.serve.
mcp Serving plus the MCP endpoint that exposes models to LLM agents.
bohb Dependencies required by the bohb search strategy and scheduler.
graph PyTorch Geometric, required by the graph-based recommenders.
all All of the above.

You can install them at any moment using the following command:

pip install "warprec[dashboard, remote-io]"

πŸ“¦ Install via Poetry

If you use Poetry for dependency management, you can easily install WarpRec and its dependencies directly from the source:

  1. Clone the repository Open your terminal and clone the WarpRec repository:

    git clone <repository_url>
    cd warprec
  2. Install the project

    poetry install
    # Or you can install all extra dependencies
    poetry install --extras all

πŸ› οΈ Development Setup (Conda)

If you want to contribute, we recommend using Conda. The environment installs WarpRec with all extra dependencies:

  1. Clone the repository Open your terminal and clone the WarpRec repository:

    git clone <repository_url>
    cd warprec
  2. Create the Conda environment Use the provided environment.yml file. It installs Python 3.12 and then WarpRec itself with all extras, so the dependency set always matches pyproject.toml.

    conda env create --file environment.yml
  3. Activate the environment:

    conda activate warprec
  4. CPU-only machines (optional)

    The environment installs the default PyTorch build, which is CUDA-enabled on Linux. On a machine without a GPU you can replace it with the smaller CPU build:

    pip install torch==2.7.* --index-url https://download.pytorch.org/whl/cpu

πŸš‚ Usage

πŸ‹οΈβ€β™‚οΈ Training a model

To train a model, use the train pipeline. Here's an example:

  1. Prepare a configuration file (e.g. config/train_config.yml) with details about the model, dataset and training parameters.
  2. Start a Ray HEAD node:
    ray start --head
  3. Run the following command:
    # Running with pip
    warprec -c config/train_config.yml -p train
    # Or with cloned repo
    python -m warprec.run -c config/train_config.yml -p train

This command starts the training process using the specified configuration file.

✏️ Design a model

To implement a custom model, WarpRec provides a dedicated design interface via the design pipeline. The recommended workflow is as follows:

  1. Prepare a configuration file (e.g. config/design_config.yml) with details about the custom models, dataset and training parameters.
  2. Run the following command:
    # Running with pip
    warprec -c config/design_config.yml -p design
    # Or with cloned repo
    python -m warprec.run -c config/design_config.yml -p design

This command initializes a lightweight training pipeline, specifically intended for rapid prototyping and debugging of custom architectures within the framework.

πŸ” Evaluate a model

To run only evaluation on a model, use the eval pipeline. Here's an example:

  1. Prepare a configuration file (e.g. config/eval_config.yml) with details about the model, dataset and training parameters.
  2. Run the following command:
    # Running with pip
    warprec -c config/eval_config.yml -p eval
    # Or with cloned repo
    python -m warprec.run -c config/eval_config.yml -p eval

This command starts the evaluation process using the specified configuration file.

πŸ›°οΈ Serve a model

A model trained with meta.save_model: true can be served directly from its checkpoint:

  1. Install the serving extra:
    pip install "warprec[serving]"
  2. Prepare a serving configuration (e.g. config/serve_config.yml) that points each endpoint at a saved .pth file.
  3. Start the server:
    # Running with pip
    warprec.serve -c config/serve_config.yml
    # Or with cloned repo
    python -m warprec.serve -c config/serve_config.yml
  4. Ask for recommendations:
    curl -X POST localhost:8000/v1/models/sasrec/recommend \
         -H "Content-Type: application/json" -d '{"user_id": 1, "k": 10}'

See the serving guide for the API, MCP, scaling and cluster deployment.

🧰 Makefile Commands

The project includes a Makefile to simplify common operations:

  • 🧹 Run linting:
    make lint
  • πŸ§‘β€πŸ”¬ Run tests:
    make test

πŸ““ Guides

The guides/ folder holds 17 Jupyter notebooks, committed with their outputs, that go through WarpRec one piece at a time on public datasets: reading, filtering and splitting data, side information, contexts, knowledge graphs and multimodal features, every dataloader, the model families, hyperparameter search, every metric family, cold start and debiased evaluation, custom models and metrics, callbacks, and serving. Each step runs through the Python API and through the configuration that does the same, so the notebooks double as worked examples of both. See guides/README.md or the Guides section of the documentation.

🀝 Contributing

We welcome contributions from the community! Whether you're fixing bugs, improving documentation, or proposing new features, your input is highly valued.

To get started:

  1. Fork the repository and create a new branch for your feature or fix.
  2. Follow the existing coding style and conventions.
  3. Make sure the code passes all checks by running make lint.
  4. Open a pull request with a clear description of your changes.

If you encounter any issues or have questions, feel free to open an issue in the Issues section of the repository.

πŸ“œ License

This project is licensed under the MIT License - see the LICENSE file for details.

πŸ“– Citation

Citation details will be provided in an upcoming release. Stay tuned!

πŸ“§ Contact

For questions or suggestions, feel free to contact us at:

Releases

Packages

Contributors

Languages