Upload a video, get a full transcript, an executive report, and a professional infographic β powered by AI.
TranscribeVideos is a fullstack application that extracts audio from video files, transcribes them using your choice of 5 STT providers, and generates structured AI summaries with dual-pass extraction plus professional infographics with charts and data tables. Export results to PDF or Word with one click.
Designed to be cost-efficient (~$0.02 for a 2-hour video with GPT-4o-mini summarization) and ADHD-friendly with visual infographics and color-coded sections.
- 5 STT providers β OpenAI Whisper, Google Cloud STT, Deepgram, Qwen/DashScope, MiniMax. Configure your own API keys per provider.
- Automatic chunking β Files over 24 MB are split into segments and transcribed in parallel.
- Multi-format support β MP4, MOV, AVI, MKV, WebM, FLV, WMV, M4V, 3GP (video) and MP3, WAV, M4A, OGG, FLAC, OPUS, AAC, WMA (audio).
- Pass 1 β Exhaustive extraction β Every fact, number, decision, example, and topic is extracted from the raw transcript.
- Pass 2 β Structured analysis β The extracted data is organized into a detailed JSON with:
- Main idea β one powerful sentence
- Topics (15β35) β each with title and detailed description
- Key insights (8β18) β with concrete data points
- Conclusions (5β12) β assertion + justification
- Action items β tasks, deadlines, owners
- Stats & facts β numbers, percentages, metrics
- Key decisions β decisions made and their rationale
- Executive summary β 8β12 sentence paragraph plus key takeaways
- AI-generated structured data rendered with Chart.js (bar/doughnut charts)
- Metric cards with trend indicators
- Data tables for stats and figures
- Topic deep-dive sections with icons and key points
- Visual timeline of content flow
- Conclusion panel with takeaways and next steps
- PDF β High-quality export of both the summary report and infographic (via html2pdf.js)
- Word (.docx) β Formatted document with cover page, stats, all sections, and tables
- HTML β Raw infographic HTML for embedding elsewhere
- Dark theme β Clean, distraction-free interface
- Live progress β Step-by-step sub-tasks with checkmarks and elapsed timer
- Transcribe-only mode β Skip summarization for faster, cheaper results
- Cost transparency β Detailed breakdown: STT cost + GPT tokens + total
- Job history β SQLite-backed persistence. Past jobs survive server restarts
- Configurable prompts β Customize the extraction, structuring, and infographic AI prompts via the Settings panel
- Settings panel β Configure API keys per provider, default models, language, and prompt overrides
- Electron wrapper β Native macOS/Windows app that bundles ffmpeg/ffprobe binaries
- No external dependencies needed for end users
ASCII mockups representing the three main result views
βββ Upload βββββββββββββββββββββββββββββββββββββββββββ
β β
β [Drag & drop your video here] β
β β
β STT Provider: [OpenAI Whisper βΌ] Model: [gpt-4o-mini βΌ] β
β β‘ Transcribe only (skip summary) β
β β
β βββ History ββββββββββββββββββββββββββββββββββββ β
β meeting-01.mp4 Done $0.74 2h ago β
β conference.mp4 Done $1.23 Yesterday β
βββ Processing βββββββββββββββββββββββββββββββββββββββ
β β
β β New Video meeting-01.mp4 00:03:42 β
β ββββββββββββββββββββββββββ 46% [transcribing] β
β β
β β Analyzing video file... β
β β Extracting audio track... β
β β Encoding to optimal format (16kHz, mono)... β
β β Audio ready: 1h 23m duration β
β β Transcribing chunk 3 of 5... β
β β Generating executive summary... β
β β Creating infographic... β
βββ Results ββββββββββββββββββββββββββββββββββββββββββ
β [Transcript] [Summary] [Infographic] [PDF] [Word] β
β β
β ββ Stats ββββββββββββββββββββββββββββββββββββββββ β
β β 24 Topics β 12 Insights β 6 Actions β ... β β
β βββββββββββββββββββββββββββββββββββββββββββββββββ β
β β
β ββ Key Takeaways ββββββββββββββββββββββββββββββββ β
β β 01 Revenue grew 27% YoY driven by APAC... β β
β β 02 Customer churn dropped to 4.2% after... β β
β βββββββββββββββββββββββββββββββββββββββββββββββββ β
β β
β ββ Topics ββββββββββββββββββββββββββββββββββββββββ β
β β 01 Market expansion strategy β 02 Q3 roadmap β β
β β 03 Hiring plan for 2026 β 04 Budget... β β
β βββββββββββββββββββββββββββββββββββββββββββββββββ β
β β
β Cost: $0.74 (Whisper: $0.72 + GPT: $0.02) β
transcribe-videos/
βββ server/ # Backend (Express + Node.js)
β βββ src/
β βββ index.js # Server entry point (dynamic port)
β βββ routes/
β β βββ jobs.js # REST API: upload, poll, export, delete
β β βββ settings.js # GET/PUT settings & provider config
β βββ services/
β βββ extractAudio.js # ffmpeg: video β MP3 extraction
β βββ media-binaries.js # ffmpeg/ffprobe binary resolution
β βββ database.js # SQLite persistence (better-sqlite3, WAL)
β βββ settings-store.js # Settings CRUD from SQLite
β βββ exportDocx.js # Word (.docx) document generation
β βββ transcribe/
β β βββ index.js # Multi-provider orchestration
β β βββ openai.js # OpenAI Whisper provider
β β βββ google.js # Google Cloud STT provider
β β βββ qwen.js # Qwen / DashScope provider
β β βββ minimax.js # MiniMax provider
β β βββ deepgram.js # Deepgram provider
β βββ summarize/
β βββ index.js # Two-pass summary orchestration
β βββ defaults.js # AI prompts (extraction, structuring, infographic)
β βββ infographic-data.js # Structured infographic data generation
βββ client/ # Frontend (Vue 3 + Vite)
β βββ src/
β β βββ App.vue # Root component + screen router
β β βββ api.js # Fetch wrapper for all endpoints
β β βββ style.css # Global styles & dark theme
β β βββ composables/
β β β βββ useJobPolling.js # Reactive polling composable
β β βββ components/
β β βββ UploadPanel.vue # Drag & drop file upload + settings
β β βββ ProgressTracker.vue # Live progress with sub-steps
β β βββ TranscriptView.vue # Full transcript display
β β βββ SummaryView.vue # Professional executive report
β β βββ InfographicView.vue # Chart.js infographic with export
β β βββ JobHistory.vue # Past transcription list
β β βββ SettingsModal.vue # Settings & provider config
β β βββ ProviderSettings.vue # API key config per STT provider
β β βββ ModelSettings.vue # Default model selection
β β βββ ProviderSelector.vue # STT provider picker
β β βββ PromptSettings.vue # Custom AI prompt editor
β βββ dist/ # Production build (served by Express)
βββ desktop/ # Electron desktop app wrapper
β βββ main.js # Electron main process
β βββ preload.js # Context bridge
β βββ electron-builder.yml # Build config for macOS/Windows
βββ scripts/
β βββ test.js # OpenAI API connectivity check
βββ .env.example # Environment variables template
βββ package.json # Root scripts (build, start, dev, desktop)
[Video upload] β POST /api/jobs
β
βΌ
ββ Step 1: Audio Extraction ββββββββββββββββββββββββββββββ
β ffmpeg: extract audio track β MP3 (32kbps, mono, 16kHz)β
β Duration detected via ffprobe β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
ββ Step 2: Transcription βββββββββββββββββββββββββββββββββ
β If MP3 > 24 MB: split into chunks β
β Each chunk β selected STT provider: β
β β’ OpenAI Whisper ($0.006/min) β
β β’ Google Cloud STT ($0.016/min) β
β β’ Deepgram ($0.0125/min) β
β β’ Qwen/DashScope ($0.004/min) β
β β’ MiniMax ($0.003/min) β
β Results concatenated into full transcript β
β Retry logic: 3 attempts per chunk β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
ββ Step 3: Summarization (skipped if transcribe-only) ββββ
β Two-pass mode (default for transcripts > 1500 words): β
β Pass 1: extractData() β exhaustive raw extraction β
β Pass 2: structureData() β structured JSON output β
β β
β Single-pass fallback for short transcripts β
β Model: GPT-4o-mini or GPT-4o β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
ββ Step 4: Infographic Generation ββββββββββββββββββββββββ
β AI generates structured JSON data: β
β β’ Hero section, metric cards, chart configs β
β β’ Data tables, topic deep-dives, timeline β
β β’ Conclusion with takeaways and next steps β
β β
β Rendered client-side with Chart.js and SVG icons β
β Fallback: legacy GPT-generated HTML β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
[Frontend polls GET /api/jobs/:id every 2.5s]
[Results displayed with 3 tabs: Transcript, Summary, Infographic]
[Export: PDF (html2pdf.js) and Word (docx)]
- Node.js 18+ (LTS recommended)
- OpenAI API key β Get one here
- ffmpeg (optional for standalone server) β the server bundles precompiled binaries for most platforms
| Provider | API Key Required | Get Key |
|---|---|---|
| OpenAI Whisper | Yes | platform.openai.com |
| Google Cloud STT | Yes | console.cloud.google.com |
| Deepgram | Yes | console.deepgram.com |
| Qwen / DashScope | Yes | dashscope.aliyun.com |
| MiniMax | Yes | minimax.io |
# macOS
brew install ffmpeg
# Ubuntu / Debian
sudo apt install ffmpeg
# Windows (with Chocolatey)
choco install ffmpeg# 1. Clone the repository
git clone https://github.com/ecr17dev/transcribe-videos.git
cd transcribe-videos
# 2. Install all dependencies
npm install
cd server && npm install && cd ..
cd client && npm install && cd ..
# 3. Configure your OpenAI API key (required)
cp .env.example .env
# Edit .env and add your key:
# OPENAI_API_KEY=sk-your-key-here
# 4. Verify connectivity
npm test
# 5. Start the app (builds client + starts server)
npm start
# Open http://localhost:6969npm run dev
# Backend: http://localhost:6969 (with --watch, auto-reload)
# Frontend: http://localhost:5173 (Vite HMR, proxies /api to backend)# Install desktop dependencies (with native module rebuild)
npm run desktop:install
# Run in development mode
npm run desktop:dev
# Build for macOS
npm run desktop:build:mac
# Build for Windows
npm run desktop:build:winEdit .env to configure the server:
# Required: OpenAI API key
OPENAI_API_KEY=sk-your-key-here
# Optional: Server port (default: 6969, falls back to next available if busy)
PORT=6969
# Optional: Vite proxy target port for development
SERVER_PORT=6969Additional STT provider API keys and settings are configured via the in-app Settings panel (gear icon). These are persisted in the SQLite database and survive server restarts.
| Setting | Description | Default |
|---|---|---|
| Default STT provider | Speech-to-text engine to use | OpenAI |
| Default summary model | GPT model for summarization | gpt-4o-mini |
| Two-pass summary | Use dual-pass extraction + structuring | Enabled |
| Default language | Prompt language preference | Spanish |
| Extraction prompt | Custom prompt for Pass 1 (data extraction) | β |
| Structuring prompt | Custom prompt for Pass 2 (JSON structuring) | β |
| Infographic prompt | Custom prompt for legacy HTML infographic | β |
| Method | Path | Description |
|---|---|---|
POST |
/api/jobs |
Upload a video and start processing |
GET |
/api/jobs |
List all transcription jobs (history) |
GET |
/api/jobs/:id |
Get job status and results |
DELETE |
/api/jobs/:id |
Delete a transcription job |
GET |
/api/jobs/:id/export/docx |
Download Word (.docx) report |
GET |
/api/jobs/:id/export/html |
Download printable HTML report |
GET |
/api/settings |
Get all settings |
PUT |
/api/settings |
Update settings |
Upload a video file as multipart/form-data.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
video |
File | Yes | β | Video/audio file (max 5 GB) |
model |
String | No | gpt-4o-mini |
GPT model for summary (gpt-4o-mini, gpt-4o) |
transcribeOnly |
String | No | false |
Set to "true" to skip summarization |
sttProvider |
String | No | openai |
STT provider: openai, google, deepgram, qwen, minimax |
sttModel |
String | No | whisper-1 |
Provider-specific model identifier |
Response: { "jobId": "uuid-string" }
Poll for job progress and retrieve results.
Response (processing):
{
"id": "uuid",
"status": "transcribing",
"step": 2,
"progress": 46,
"stepDetail": "Chunk 3 of 5...",
"subSteps": [
{ "text": "Audio ready: 1h 23m", "done": true },
{ "text": "Transcribing chunk 3 of 5", "done": false }
],
"transcript": null,
"summary": null,
"infographic": null,
"infographicData": null,
"detailedSummary": null,
"cost": null,
"error": null,
"videoName": "meeting.mp4",
"model": "gpt-4o-mini",
"sttProvider": "openai",
"sttModel": "whisper-1",
"startedAt": 1717536000000
}When status === "done", additional fields:
transcriptβ Full transcription textsummaryβ Normalized summary object (topics, insights, conclusions, etc.)detailedSummaryβ Raw JSON from AI (withdetailed_breakdownandexecutive_summary)infographicβ Legacy GPT-generated HTML (fallback)infographicDataβ Structured JSON for the Chart.js infographiccostβ Breakdown:{ stt, gpt, total, durationMinutes, inputTokens, outputTokens, twoPass }
Status values: extracting β transcribing β summarizing β done | error
| Provider | Cost per minute | 1-hour video | 2-hour video |
|---|---|---|---|
| MiniMax | $0.003 | ~$0.18 | ~$0.36 |
| Qwen / DashScope | $0.004 | ~$0.24 | ~$0.48 |
| OpenAI Whisper | $0.006 | ~$0.36 | ~$0.72 |
| Deepgram | $0.0125 | ~$0.75 | ~$1.50 |
| Google Cloud STT | $0.016 | ~$0.96 | ~$1.92 |
| Model | Input tokens (1K) | Output tokens (1K) | Typical cost for 2h video |
|---|---|---|---|
| GPT-4o Mini | $0.00015 | $0.0006 | ~$0.02 |
| GPT-4o | $0.0025 | $0.01 | ~$0.25 |
| Model | Typical cost |
|---|---|
| GPT-4o Mini | ~$0.01 |
| Provider | Model | Mode | 1-hour | 2-hour |
|---|---|---|---|---|
| OpenAI Whisper | GPT-4o Mini | Full | ~$0.39 | ~$0.77 |
| OpenAI Whisper | GPT-4o Mini | Transcribe-only | ~$0.36 | ~$0.72 |
| MiniMax | GPT-4o Mini | Full | ~$0.21 | ~$0.39 |
| Qwen | GPT-4o Mini | Full | ~$0.27 | ~$0.51 |
Costs are estimates based on API pricing. Actual costs depend on audio duration, transcript length, and token usage. The app displays a detailed breakdown after each job completes, including per-provider STT cost, GPT input/output token counts, and whether two-pass summarization was used.
| Layer | Technology |
|---|---|
| Runtime | Node.js 18+ |
| Backend | Express 4 |
| Frontend | Vue 3 (Composition API, <script setup>) |
| Build | Vite 6 |
| STT Providers | OpenAI Whisper, Google Cloud STT, Deepgram, Qwen/DashScope, MiniMax |
| AI Summarization | OpenAI GPT-4o Mini / GPT-4o |
| Charts | Chart.js 4 |
| PDF Export | html2pdf.js |
| Word Export | docx |
| Icons | Tabler Icons |
| Audio Processing | ffmpeg (fluent-ffmpeg) |
| Database | SQLite (better-sqlite3, WAL mode) |
| File Upload | Multer |
| Desktop | Electron 35 |
Contributions are welcome and appreciated. Here's how to get started:
- Fork the repository
- Create a feature branch:
git checkout -b feature/amazing-feature - Make your changes and commit:
git commit -m 'feat: add amazing feature' - Push to your fork:
git push origin feature/amazing-feature - Open a Pull Request
Please ensure your code follows the existing style conventions and includes relevant tests when applicable.
MIT Β© Esteban CortΓ©s
