A comprehensive web-based AI model benchmarking system for evaluating and comparing AI models across multiple providers and problem sets. | δΈζReadMe
ModelForge is designed to help developers, researchers, and organizations systematically evaluate AI models across different providers (OpenAI, Anthropic, Google Gemini, etc.) using standardized problem sets. It provides a complete benchmarking platform with automated judging, manual review capabilities, and comprehensive analytics.
- OpenAI-compatible endpoints (OpenAI, OpenRouter, local vLLM/llama.cpp)
- Anthropic Claude integration
- Google Gemini REST API
- Custom HTTP adapters for experimental providers
- Text-based problems with exact/regex/fuzzy matching
- HTML/CSS/JS tasks with DOM-based evaluation
- Automated LLM judging using neutral models
- Manual review override for complex cases
- N-way battle mode for pairwise model comparisons
- Windows 11-inspired dark theme with glass effects
- Real-time dashboards with interactive charts
- Live streaming of model responses during runs
- Responsive design for desktop and mobile
- Model performance rankings with ELO-like ratings
- Accuracy and latency distributions
- Cost analysis across providers
- Problem difficulty analysis
- Win rate matrices for battle mode
- Encrypted API keys using AES-GCM
- Secure HTML sandbox with CSP and iframe isolation
- No API key exposure to frontend
- Rate limiting and CORS protection
- Node.js 18.18.0 or higher
- npm 9.0.0 or higher
# Clone the repository
git clone <repository-url>
cd model-forge
# Install all dependencies
npm run install:all
# Start development servers (API + Web)
npm run startThe application will be available at:
- Web UI: http://localhost:5175
- API: http://localhost:5174
# Install shared packages
npm run install:shared
# Install API dependencies
npm run install:api
# Install Web dependencies
npm run install:web
# Start individual services
npm run start:api # API server only
npm run start:web # Web UI only-
Navigate to Providers & Models in the sidebar
-
Click Add Provider
-
Configure:
- Name: Display name (e.g., "OpenAI GPT-4")
- Adapter: Provider type (OpenAI, Anthropic, Gemini, Custom)
- Base URL: API endpoint (e.g.,
https://api.openai.com/v1) - API Key: Your provider API key (encrypted at rest)
- Default Model: Primary model for this provider
-
Click Test Connection to validate
-
Save the provider
- Select a provider from the list
- Click Add Model
- Configure:
- Label: Display name (e.g., "GPT-4 Turbo")
- Model ID: Provider-specific model identifier
- Settings: Optional model parameters (temperature, max_tokens, etc.)
- Navigate to Problem Sets
- Click New Problem Set
- Enter:
- Name: Descriptive title
- Description: Optional detailed description
- Save and add problems
-
Select a problem set
-
Click Add Problem
-
Choose problem type:
- Text: Natural language tasks
- HTML: Web development tasks
-
Configure:
- Prompt: The task description
- Expected Answer: For text problems
- HTML Assets: For HTML problems (HTML/CSS/JS)
- Scoring Rules: Custom evaluation criteria
-
Navigate to Runs
-
Click New Run
-
Configure:
- Name: Optional run identifier
- Problem Set: Select from available sets
- Models: Choose 2-8 models to compare
- Judge Model: Select a model for automated judging
- Streaming: Enable real-time response viewing
-
Click Create Run
- From the runs list, click Start on your new run
- Monitor progress in real-time:
- Live tokens streaming for each model
- Problem-by-problem status updates
- Completion percentages for each model
- Navigate to Dashboard
- View:
- Overall accuracy across models
- Model performance rankings
- Problem difficulty analysis
- Cost and latency metrics
- Click on any completed run
- View:
- Problem Γ Model matrix with verdicts
- Individual responses and judgments
- Manual override options for disputed results
- Navigate to Review
- For each HTML task:
- View live sandbox rendering
- Compare expected vs actual output
- Override automated judgments if needed
- Navigate to Battle (coming soon)
- Select models for pairwise comparison
- View win rate matrices and ELO ratings
- Analyze statistical significance of results
model-forge/
βββ apps/
β βββ api/ # Fastify TypeScript API
β β βββ src/
β β β βββ server.ts # Main server with 2000+ lines
β β βββ package.json
β βββ web/ # React TypeScript frontend
β βββ src/
β β βββ app/ # Routes and layouts
β β βββ features/ # Domain modules
β β βββ components/ # Design system
β β βββ lib/ # API and utilities
β βββ package.json
βββ packages/
β βββ shared/ # Shared types and utilities
βββ package.json # Root workspace configuration
- Frontend: React 18, TypeScript, Vite, Tailwind CSS
- Backend: Fastify, TypeScript, SQLite (better-sqlite3)
- Database: SQLite with encrypted API keys
- Charts: Apache ECharts for analytics
- Animation: Framer Motion for smooth transitions
- Forms: React Hook Form with Zod validation
# Development
npm run dev # Start all services
npm run start # Alias for dev
npm run start:api # API server only
npm run start:web # Web UI only
# Building
npm run build # Build all packages
npm run build:api # Build API only
npm run build:web # Build web only
# Quality
npm run typecheck # Type checking across all packages
npm run lint # Lint all packages
npm run format # Format code with PrettierCreate .env files in respective directories:
apps/api/.env
PORT=5174
ENCRYPTION_KEY=your-32-char-encryption-keyapps/web/.env
VITE_API_URL=http://localhost:5174- Location: Auto-created at
apps/api/apps/api/var/data.sqlite - Schema: Auto-created on server start
- Encryption: API keys encrypted with AES-GCM
- Cleanup: Use
npm run clean:dbto remove database files for a fresh start - Backup: SQLite file can be copied for backup
- Ensure database directory is writable
- Check if another instance is running
- Use
npm run clean:dbto reset database (will lose data)
- Verify provider endpoints are accessible
- Check API key permissions
- Use the Test Connection feature before saving
- Ensure API and web are running on expected ports
- Check
.envconfiguration matches actual URLs
- Concurrent model evaluation: 2-8+ models simultaneously
- Problem sets: Unlimited problems per set
- Real-time streaming: Live token updates during runs
- Scoring accuracy: Automated + manual override
- Cost tracking: Per-model and per-run cost analysis
- Memory: 512MB+ RAM for basic usage
- Storage: 100MB+ for database and logs
- Network: Stable internet for API calls
npm run build- API: Render, Fly.io, or Railway
- Web: Vercel, Netlify, or GitHub Pages
- Database: SQLite with Litestream for persistence
- Fork the repository
- Create a feature branch:
git checkout -b feature/amazing-feature - Commit changes:
git commit -m 'Add amazing feature' - Push to branch:
git push origin feature/amazing-feature - Open a Pull Request
- Windows 11 Design System for visual inspiration
- Fastify for high-performance API framework
- React ecosystem for modern web development
- Better SQLite3 for reliable database operations
Built with β€οΈ for the AI benchmarking community
