Self-hosted Slippi replay archive system for lunarmelee.com.
Crawls ~20TB of .slp replay files, indexes metadata into MongoDB, and serves an API for searching/filtering replays and requesting download bundles.
┌─────────────────────────────────────────────────────────┐
│ Home Machine │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────────────┐ │
│ │ Crawler │──▶│ MongoDB │◀──│ Express API │◀───┼──── Cloudflare Tunnel
│ └──────────┘ └──────────┘ │ /api/replays │ │ │
│ │ /api/jobs │ │ │
│ ┌──────────┐ │ /api/stats │ │ lunarmelee.com
│ │ Job │◀─────────────────│ /api/jobs/:id/dl │ │ (Next.js on Linode)
│ │ Worker │ └──────────────────┘ │
│ └──────────┘ │
│ │ │
│ ▼ │
│ ┌──────────┐ │
│ │ Bundles │ (.zip files served for download) │
│ │ /data/ │ │
│ └──────────┘ │
└─────────────────────────────────────────────────────────┘
- Crawler — Walks the
.slpdirectory tree, parses each file with@slippi/slippi-js, stores metadata in MongoDB - API — Express server exposing replay search, job creation, job status, and bundle downloads
- Job Worker — Polls for pending download jobs, bundles matching replays into
.zipfiles - Bundle Cache — Completed bundles are kept for a configurable TTL, then cleaned up automatically
# Install dependencies
npm install
# Copy env and configure
cp .env.example .env
# Edit .env with your MongoDB URI, SLP root directory, etc.
# Development
npm run dev
# Build & run
npm run build
npm startIndex your replay files:
# Uses SLP_ROOT_DIR from .env
npm run crawl
# Or specify a directory
npm run crawl -- /path/to/slp/filesThe crawler skips files that are already indexed (by file path) and never follows symlinks (the archive keeps ~1M flat-path links for replay_archiver; indexing them would duplicate replays). Safe to re-run.
After every crawl, rebuild the player autocomplete summary, which the crawler does not maintain:
npm run build-playersThe "Download Full DB" button serves one pre-built archive,
b2:lm-replays/archive/lunar_db_full.zip: every real replay's .slpz (raw .slp
where slpz can't compress it). It is a snapshot; the site shows its date, size and
replay count. Rebuild it on the worker after large imports:
nohup scripts/full-db/rebuild.sh > /dev/null 2>&1 & # build → verify → upload → register
tail -f ~/Projects/worker/shared_folder_2/full_db/rebuild.logThe build writes to shared_folder_2 (a different disk from the archive) and takes
a few hours; the upload takes days. The old archive keeps serving until the upload
completes. Rerunning after a failure reuses an unpublished build. The last step,
npm run register-full-db -- --size BYTES --replays N --snapshot ISO_DATE, updates
the pinned bundle record the site reads.
extract-stats fills the gameStats collection (one document per replay) and
writes per-game event files. Extraction is split into independently versioned
extractors (EXTRACTORS in src/services/gameStats.ts): core (context,
result, per-player stats; conversions, deaths), clipper (Clipper's combos,
edgeguards, phantoms, early quit-outs), identity (content hash, cross-recording
fingerprint, Gecko codes), position (stage position/posture) and techLedge
(tech, getup and ledge options). Each writes only its own fields plus
extractors.<name> (its version), and its events to
<detail-dir>/<name>/v<version>/<shard>.jsonl.gz.
A run names the extractors it computes and is fixed to their versions
(statsRuns). The first run computes everything; later runs add a new extractor
or recompute one whose version was bumped, leaving the rest untouched. Plan a run
again after a crawl to extend it to the new replays.
Work is split into shards (replay-ID ranges in statsShards) claimed in random
order with expiring leases, so any number of machines can run it against the
same MongoDB and detail directory. A shard is done only once committed; a killed
runner's shard is reclaimed when its lease expires.
npm run extract-stats -- --run main --plan [--extractors core,clipper] [--shard-size 5000]
npm run extract-stats -- --run main --detail-dir DIR [--workers N] [--max-shards N]
npm run extract-stats -- --run main --status
npm run extract-stats -- --runs
STATS_NAMESPACE=pilot npm run extract-stats -- ... # trial run in *_pilot collectionsAnother machine needs the replay files (read only), the detail directory
(read-write), slpz, and MongoDB through an SSH tunnel (don't expose MongoDB on the
LAN): scripts/stats/run-extractor.sh opens the tunnel and starts a runner from
~/.config/lunar-stats/env. Shards that fail three times stay failed in
--status. After a run, npm run backfill-match-info -- --apply copies match info
onto replays crawled before it was recorded at crawl time.
Player profiles (playerStats, served by GET /api/players/:code/profile) are
rebuilt from gameStats with npm run build-player-stats (about 2.5 minutes for
the whole archive). Run it after a stats run and after extending a run to new
replays. Profiles cover human 1v1 games per connect code; Slippi user IDs and the
alternate codes they link are stored but not served publicly.
Search/filter replays. Query params:
| Param | Description |
|---|---|
connectCode |
Player connect code (e.g. MATT#123) |
characterId |
Character ID number |
stageId |
Stage ID number |
startDate |
ISO date string, lower bound |
endDate |
ISO date string, upper bound |
page |
Page number (default 1) |
limit |
Results per page (default 50, max 200) |
Get a single replay by ID.
Create a download job. Body (JSON):
{
"connectCode": "MATT#123",
"characterId": 2,
"stageId": 31,
"startDate": "2023-01-01",
"endDate": "2024-01-01"
}All fields optional. Returns { jobId, status }.
Check job status. Returns { jobId, status, bundleSize, error, createdAt, completedAt }.
Download the completed bundle (.zip).
Returns { replays: count, jobs: { pending, processing, completed, failed } }.
Returns { ok: true }.
| Variable | Default | Description |
|---|---|---|
MONGODB_URI |
mongodb://localhost:27017/lm-database |
MongoDB connection string |
PORT |
3000 |
API server port |
SLP_ROOT_DIR |
/data/slp |
Root directory of .slp files |
LUNAR_SERVICE_KEY |
empty | Shared secret identifying the website (same value in the website's env). When set, only the website may forward X-Visitor-Ip (per-visitor rate limits) or send X-Client-Id; other callers' identity headers are dropped. Deploy the website with the key before setting it here. |
src/
config.ts — Environment config
db.ts — MongoDB connection
index.ts — Express app + worker startup
models/
Replay.ts — Replay metadata schema
Job.ts — Download job schema
routes/
replays.ts — Replay search/query endpoints
jobs.ts — Job CRUD + download endpoints
stats.ts — Stats endpoint
services/
slpParser.ts — Parse .slp files via @slippi/slippi-js
crawler.ts — Directory walker + batch indexer
bundler.ts — Zip bundler + cleanup
workers/
jobWorker.ts — Job queue processor
scripts/
crawl.ts — CLI entry point for crawling
The API is exposed publicly via a Cloudflare Tunnel at api.lunarmelee.com. Everything you need to deploy on a fresh machine is in the deploy/ folder.
- Node.js (v20+)
- MongoDB (v8.0) — installed and running as
mongodsystemd service - cloudflared — install instructions
- Tunnel credentials — the
208535eb-4007-4e92-9d8d-4e3ab1c530c5.jsonfile (keep this secret, copy from previous machine)
# 1. Clone the repo
git clone https://github.com/madenney/lm-database.git
cd lm-database
# 2. Install dependencies and build
npm install
npm run build
# 3. Configure environment
cp .env.example .env
# Edit .env — fill in paths, R2 keys, JWT secret, etc.
# Generate a JWT secret: openssl rand -hex 32
# 4. Copy tunnel credentials into place
# Get the credentials JSON from the old machine (~/.cloudflared/*.json)
# and put it at: ~/.cloudflared/208535eb-4007-4e92-9d8d-4e3ab1c530c5.json
# 5. Make sure MongoDB is running
sudo systemctl start mongod
sudo systemctl enable mongod
# 6. Run the setup script (installs systemd services, starts everything)
sudo bash deploy/setup.sh- Copies tunnel credentials to
/etc/cloudflared/credentials.json - Copies tunnel config to
/etc/cloudflared/config.yml - Installs two systemd services:
lm-database-api— the Express API serverlm-database-tunnel— the Cloudflare tunnel
- Enables and starts both services (they auto-start on boot)
# Check status
sudo systemctl status lm-database-api
sudo systemctl status lm-database-tunnel
# View logs (live)
sudo journalctl -u lm-database-api -f
sudo journalctl -u lm-database-tunnel -f
# Restart after code changes
npm run build
sudo systemctl restart lm-database-api
# Stop everything
sudo systemctl stop lm-database-api lm-database-tunnelThis should only be necessary if you've lost the credentials file.
cloudflared tunnel login # opens browser, pick lunarmelee.com
cloudflared tunnel create lm-database # creates new tunnel + credentials JSON
cloudflared tunnel route dns lm-database api.lunarmelee.com # set DNS
# Then update the tunnel ID in deploy/cloudflared.yml and re-run setupdeploy/
├── cloudflared.yml # Tunnel config (api.lunarmelee.com → localhost:3002)
├── lm-database-api.service # systemd unit for Express API
├── lm-database-tunnel.service # systemd unit for cloudflared
└── setup.sh # Installs everything, run with sudo
- Database migration — Schema is designed to be portable. If MongoDB can't handle hundreds of millions of records, migrate to Postgres.
- Concurrency — Job worker currently processes one job at a time. Can be scaled up.