Skip to content

Latest commit

Β 

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

TranscribeVideos

Upload a video, get a full transcript, an executive report, and a professional infographic β€” powered by AI.

Node.js Vue.js License


Overview

TranscribeVideos is a fullstack application that extracts audio from video files, transcribes them using your choice of 5 STT providers, and generates structured AI summaries with dual-pass extraction plus professional infographics with charts and data tables. Export results to PDF or Word with one click.

Designed to be cost-efficient (~$0.02 for a 2-hour video with GPT-4o-mini summarization) and ADHD-friendly with visual infographics and color-coded sections.


Demo

TranscribeVideos Demo


Features

Transcription

  • 5 STT providers β€” OpenAI Whisper, Google Cloud STT, Deepgram, Qwen/DashScope, MiniMax. Configure your own API keys per provider.
  • Automatic chunking β€” Files over 24 MB are split into segments and transcribed in parallel.
  • Multi-format support β€” MP4, MOV, AVI, MKV, WebM, FLV, WMV, M4V, 3GP (video) and MP3, WAV, M4A, OGG, FLAC, OPUS, AAC, WMA (audio).

AI Summary (Two-Pass Extraction)

  • Pass 1 β€” Exhaustive extraction β€” Every fact, number, decision, example, and topic is extracted from the raw transcript.
  • Pass 2 β€” Structured analysis β€” The extracted data is organized into a detailed JSON with:
    • Main idea β€” one powerful sentence
    • Topics (15–35) β€” each with title and detailed description
    • Key insights (8–18) β€” with concrete data points
    • Conclusions (5–12) β€” assertion + justification
    • Action items β€” tasks, deadlines, owners
    • Stats & facts β€” numbers, percentages, metrics
    • Key decisions β€” decisions made and their rationale
    • Executive summary β€” 8–12 sentence paragraph plus key takeaways

Professional Infographic

  • AI-generated structured data rendered with Chart.js (bar/doughnut charts)
  • Metric cards with trend indicators
  • Data tables for stats and figures
  • Topic deep-dive sections with icons and key points
  • Visual timeline of content flow
  • Conclusion panel with takeaways and next steps

Export

  • PDF β€” High-quality export of both the summary report and infographic (via html2pdf.js)
  • Word (.docx) β€” Formatted document with cover page, stats, all sections, and tables
  • HTML β€” Raw infographic HTML for embedding elsewhere

UX

  • Dark theme β€” Clean, distraction-free interface
  • Live progress β€” Step-by-step sub-tasks with checkmarks and elapsed timer
  • Transcribe-only mode β€” Skip summarization for faster, cheaper results
  • Cost transparency β€” Detailed breakdown: STT cost + GPT tokens + total
  • Job history β€” SQLite-backed persistence. Past jobs survive server restarts
  • Configurable prompts β€” Customize the extraction, structuring, and infographic AI prompts via the Settings panel
  • Settings panel β€” Configure API keys per provider, default models, language, and prompt overrides

Desktop App

  • Electron wrapper β€” Native macOS/Windows app that bundles ffmpeg/ffprobe binaries
  • No external dependencies needed for end users

Screenshots

ASCII mockups representing the three main result views

 ─── Upload ───────────────────────────────────────────
β”‚                                                        β”‚
β”‚   [Drag & drop your video here]                        β”‚
β”‚                                                        β”‚
β”‚   STT Provider: [OpenAI Whisper β–Ό]  Model: [gpt-4o-mini β–Ό] β”‚
β”‚   β–‘ Transcribe only (skip summary)                     β”‚
β”‚                                                        β”‚
β”‚   ─── History ────────────────────────────────────    β”‚
β”‚   meeting-01.mp4       Done    $0.74    2h ago         β”‚
β”‚   conference.mp4       Done    $1.23    Yesterday      β”‚
 ─── Processing ───────────────────────────────────────
β”‚                                                        β”‚
β”‚   ← New Video    meeting-01.mp4        00:03:42        β”‚
β”‚   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘  46%  [transcribing]     β”‚
β”‚                                                        β”‚
β”‚   βœ“ Analyzing video file...                            β”‚
β”‚   βœ“ Extracting audio track...                          β”‚
β”‚   βœ“ Encoding to optimal format (16kHz, mono)...        β”‚
β”‚   βœ“ Audio ready: 1h 23m duration                       β”‚
β”‚   ● Transcribing chunk 3 of 5...                       β”‚
β”‚   β—‹ Generating executive summary...                    β”‚
β”‚   β—‹ Creating infographic...                            β”‚
 ─── Results ──────────────────────────────────────────
β”‚  [Transcript]  [Summary]  [Infographic]      [PDF] [Word] β”‚
β”‚                                                        β”‚
β”‚  β”Œβ”€ Stats ───────────────────────────────────────┐   β”‚
β”‚  β”‚  24 Topics  β”‚  12 Insights  β”‚  6 Actions  β”‚ ... β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                                                        β”‚
β”‚  β”Œβ”€ Key Takeaways ───────────────────────────────┐   β”‚
β”‚  β”‚  01  Revenue grew 27% YoY driven by APAC...   β”‚   β”‚
β”‚  β”‚  02  Customer churn dropped to 4.2% after...   β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                                                        β”‚
β”‚  β”Œβ”€ Topics ───────────────────────────────────────┐  β”‚
β”‚  β”‚  01 Market expansion strategy  β”‚  02 Q3 roadmap β”‚  β”‚
β”‚  β”‚  03 Hiring plan for 2026      β”‚  04 Budget...   β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                                                        β”‚
β”‚  Cost: $0.74 (Whisper: $0.72 + GPT: $0.02)            β”‚

Architecture

transcribe-videos/
β”œβ”€β”€ server/                              # Backend (Express + Node.js)
β”‚   └── src/
β”‚       β”œβ”€β”€ index.js                     # Server entry point (dynamic port)
β”‚       β”œβ”€β”€ routes/
β”‚       β”‚   β”œβ”€β”€ jobs.js                  # REST API: upload, poll, export, delete
β”‚       β”‚   └── settings.js             # GET/PUT settings & provider config
β”‚       └── services/
β”‚           β”œβ”€β”€ extractAudio.js          # ffmpeg: video β†’ MP3 extraction
β”‚           β”œβ”€β”€ media-binaries.js        # ffmpeg/ffprobe binary resolution
β”‚           β”œβ”€β”€ database.js             # SQLite persistence (better-sqlite3, WAL)
β”‚           β”œβ”€β”€ settings-store.js       # Settings CRUD from SQLite
β”‚           β”œβ”€β”€ exportDocx.js           # Word (.docx) document generation
β”‚           β”œβ”€β”€ transcribe/
β”‚           β”‚   β”œβ”€β”€ index.js            # Multi-provider orchestration
β”‚           β”‚   β”œβ”€β”€ openai.js           # OpenAI Whisper provider
β”‚           β”‚   β”œβ”€β”€ google.js           # Google Cloud STT provider
β”‚           β”‚   β”œβ”€β”€ qwen.js             # Qwen / DashScope provider
β”‚           β”‚   β”œβ”€β”€ minimax.js          # MiniMax provider
β”‚           β”‚   └── deepgram.js         # Deepgram provider
β”‚           └── summarize/
β”‚               β”œβ”€β”€ index.js            # Two-pass summary orchestration
β”‚               β”œβ”€β”€ defaults.js         # AI prompts (extraction, structuring, infographic)
β”‚               └── infographic-data.js # Structured infographic data generation
β”œβ”€β”€ client/                              # Frontend (Vue 3 + Vite)
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ App.vue                     # Root component + screen router
β”‚   β”‚   β”œβ”€β”€ api.js                      # Fetch wrapper for all endpoints
β”‚   β”‚   β”œβ”€β”€ style.css                   # Global styles & dark theme
β”‚   β”‚   β”œβ”€β”€ composables/
β”‚   β”‚   β”‚   └── useJobPolling.js        # Reactive polling composable
β”‚   β”‚   └── components/
β”‚   β”‚       β”œβ”€β”€ UploadPanel.vue         # Drag & drop file upload + settings
β”‚   β”‚       β”œβ”€β”€ ProgressTracker.vue     # Live progress with sub-steps
β”‚   β”‚       β”œβ”€β”€ TranscriptView.vue      # Full transcript display
β”‚   β”‚       β”œβ”€β”€ SummaryView.vue         # Professional executive report
β”‚   β”‚       β”œβ”€β”€ InfographicView.vue     # Chart.js infographic with export
β”‚   β”‚       β”œβ”€β”€ JobHistory.vue          # Past transcription list
β”‚   β”‚       β”œβ”€β”€ SettingsModal.vue       # Settings & provider config
β”‚   β”‚       β”œβ”€β”€ ProviderSettings.vue    # API key config per STT provider
β”‚   β”‚       β”œβ”€β”€ ModelSettings.vue       # Default model selection
β”‚   β”‚       β”œβ”€β”€ ProviderSelector.vue    # STT provider picker
β”‚   β”‚       └── PromptSettings.vue      # Custom AI prompt editor
β”‚   └── dist/                           # Production build (served by Express)
β”œβ”€β”€ desktop/                             # Electron desktop app wrapper
β”‚   β”œβ”€β”€ main.js                         # Electron main process
β”‚   β”œβ”€β”€ preload.js                      # Context bridge
β”‚   └── electron-builder.yml           # Build config for macOS/Windows
β”œβ”€β”€ scripts/
β”‚   └── test.js                         # OpenAI API connectivity check
β”œβ”€β”€ .env.example                        # Environment variables template
└── package.json                        # Root scripts (build, start, dev, desktop)

Data Flow

[Video upload] β†’ POST /api/jobs
       β”‚
       β–Ό
β”Œβ”€ Step 1: Audio Extraction ─────────────────────────────┐
β”‚  ffmpeg: extract audio track β†’ MP3 (32kbps, mono, 16kHz)β”‚
β”‚  Duration detected via ffprobe                          β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚
       β–Ό
β”Œβ”€ Step 2: Transcription ────────────────────────────────┐
β”‚  If MP3 > 24 MB: split into chunks                     β”‚
β”‚  Each chunk β†’ selected STT provider:                   β”‚
β”‚    β€’ OpenAI Whisper ($0.006/min)                       β”‚
β”‚    β€’ Google Cloud STT ($0.016/min)                     β”‚
β”‚    β€’ Deepgram ($0.0125/min)                            β”‚
β”‚    β€’ Qwen/DashScope ($0.004/min)                       β”‚
β”‚    β€’ MiniMax ($0.003/min)                              β”‚
β”‚  Results concatenated into full transcript             β”‚
β”‚  Retry logic: 3 attempts per chunk                     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚
       β–Ό
β”Œβ”€ Step 3: Summarization (skipped if transcribe-only) ───┐
β”‚  Two-pass mode (default for transcripts > 1500 words): β”‚
β”‚    Pass 1: extractData() β€” exhaustive raw extraction   β”‚
β”‚    Pass 2: structureData() β€” structured JSON output    β”‚
β”‚                                                        β”‚
β”‚  Single-pass fallback for short transcripts            β”‚
β”‚  Model: GPT-4o-mini or GPT-4o                          β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚
       β–Ό
β”Œβ”€ Step 4: Infographic Generation ───────────────────────┐
β”‚  AI generates structured JSON data:                    β”‚
β”‚    β€’ Hero section, metric cards, chart configs         β”‚
β”‚    β€’ Data tables, topic deep-dives, timeline           β”‚
β”‚    β€’ Conclusion with takeaways and next steps          β”‚
β”‚                                                        β”‚
β”‚  Rendered client-side with Chart.js and SVG icons      β”‚
β”‚  Fallback: legacy GPT-generated HTML                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚
       β–Ό
   [Frontend polls GET /api/jobs/:id every 2.5s]
   [Results displayed with 3 tabs: Transcript, Summary, Infographic]
   [Export: PDF (html2pdf.js) and Word (docx)]

Prerequisites

  • Node.js 18+ (LTS recommended)
  • OpenAI API key β€” Get one here
  • ffmpeg (optional for standalone server) β€” the server bundles precompiled binaries for most platforms

Optional: Additional STT Providers

Provider API Key Required Get Key
OpenAI Whisper Yes platform.openai.com
Google Cloud STT Yes console.cloud.google.com
Deepgram Yes console.deepgram.com
Qwen / DashScope Yes dashscope.aliyun.com
MiniMax Yes minimax.io

Install ffmpeg (only if bundled binaries fail)

# macOS
brew install ffmpeg

# Ubuntu / Debian
sudo apt install ffmpeg

# Windows (with Chocolatey)
choco install ffmpeg

Quick Start

# 1. Clone the repository
git clone https://github.com/ecr17dev/transcribe-videos.git
cd transcribe-videos

# 2. Install all dependencies
npm install
cd server && npm install && cd ..
cd client && npm install && cd ..

# 3. Configure your OpenAI API key (required)
cp .env.example .env
# Edit .env and add your key:
#   OPENAI_API_KEY=sk-your-key-here

# 4. Verify connectivity
npm test

# 5. Start the app (builds client + starts server)
npm start
# Open http://localhost:6969

Development Mode

npm run dev
# Backend:  http://localhost:6969 (with --watch, auto-reload)
# Frontend: http://localhost:5173 (Vite HMR, proxies /api to backend)

Desktop App

# Install desktop dependencies (with native module rebuild)
npm run desktop:install

# Run in development mode
npm run desktop:dev

# Build for macOS
npm run desktop:build:mac

# Build for Windows
npm run desktop:build:win

Configuration

Edit .env to configure the server:

# Required: OpenAI API key
OPENAI_API_KEY=sk-your-key-here

# Optional: Server port (default: 6969, falls back to next available if busy)
PORT=6969

# Optional: Vite proxy target port for development
SERVER_PORT=6969

Additional STT provider API keys and settings are configured via the in-app Settings panel (gear icon). These are persisted in the SQLite database and survive server restarts.

Configurable Settings

Setting Description Default
Default STT provider Speech-to-text engine to use OpenAI
Default summary model GPT model for summarization gpt-4o-mini
Two-pass summary Use dual-pass extraction + structuring Enabled
Default language Prompt language preference Spanish
Extraction prompt Custom prompt for Pass 1 (data extraction) β€”
Structuring prompt Custom prompt for Pass 2 (JSON structuring) β€”
Infographic prompt Custom prompt for legacy HTML infographic β€”

API Reference

Endpoints

Method Path Description
POST /api/jobs Upload a video and start processing
GET /api/jobs List all transcription jobs (history)
GET /api/jobs/:id Get job status and results
DELETE /api/jobs/:id Delete a transcription job
GET /api/jobs/:id/export/docx Download Word (.docx) report
GET /api/jobs/:id/export/html Download printable HTML report
GET /api/settings Get all settings
PUT /api/settings Update settings

POST /api/jobs

Upload a video file as multipart/form-data.

Field Type Required Default Description
video File Yes β€” Video/audio file (max 5 GB)
model String No gpt-4o-mini GPT model for summary (gpt-4o-mini, gpt-4o)
transcribeOnly String No false Set to "true" to skip summarization
sttProvider String No openai STT provider: openai, google, deepgram, qwen, minimax
sttModel String No whisper-1 Provider-specific model identifier

Response: { "jobId": "uuid-string" }

GET /api/jobs/:id

Poll for job progress and retrieve results.

Response (processing):

{
  "id": "uuid",
  "status": "transcribing",
  "step": 2,
  "progress": 46,
  "stepDetail": "Chunk 3 of 5...",
  "subSteps": [
    { "text": "Audio ready: 1h 23m", "done": true },
    { "text": "Transcribing chunk 3 of 5", "done": false }
  ],
  "transcript": null,
  "summary": null,
  "infographic": null,
  "infographicData": null,
  "detailedSummary": null,
  "cost": null,
  "error": null,
  "videoName": "meeting.mp4",
  "model": "gpt-4o-mini",
  "sttProvider": "openai",
  "sttModel": "whisper-1",
  "startedAt": 1717536000000
}

When status === "done", additional fields:

  • transcript β€” Full transcription text
  • summary β€” Normalized summary object (topics, insights, conclusions, etc.)
  • detailedSummary β€” Raw JSON from AI (with detailed_breakdown and executive_summary)
  • infographic β€” Legacy GPT-generated HTML (fallback)
  • infographicData β€” Structured JSON for the Chart.js infographic
  • cost β€” Breakdown: { stt, gpt, total, durationMinutes, inputTokens, outputTokens, twoPass }

Status values: extracting β†’ transcribing β†’ summarizing β†’ done | error


Cost Estimation

STT (Speech-to-Text)

Provider Cost per minute 1-hour video 2-hour video
MiniMax $0.003 ~$0.18 ~$0.36
Qwen / DashScope $0.004 ~$0.24 ~$0.48
OpenAI Whisper $0.006 ~$0.36 ~$0.72
Deepgram $0.0125 ~$0.75 ~$1.50
Google Cloud STT $0.016 ~$0.96 ~$1.92

Summarization (GPT)

Model Input tokens (1K) Output tokens (1K) Typical cost for 2h video
GPT-4o Mini $0.00015 $0.0006 ~$0.02
GPT-4o $0.0025 $0.01 ~$0.25

Infographic

Model Typical cost
GPT-4o Mini ~$0.01

Total (examples)

Provider Model Mode 1-hour 2-hour
OpenAI Whisper GPT-4o Mini Full ~$0.39 ~$0.77
OpenAI Whisper GPT-4o Mini Transcribe-only ~$0.36 ~$0.72
MiniMax GPT-4o Mini Full ~$0.21 ~$0.39
Qwen GPT-4o Mini Full ~$0.27 ~$0.51

Costs are estimates based on API pricing. Actual costs depend on audio duration, transcript length, and token usage. The app displays a detailed breakdown after each job completes, including per-provider STT cost, GPT input/output token counts, and whether two-pass summarization was used.


Tech Stack

Layer Technology
Runtime Node.js 18+
Backend Express 4
Frontend Vue 3 (Composition API, <script setup>)
Build Vite 6
STT Providers OpenAI Whisper, Google Cloud STT, Deepgram, Qwen/DashScope, MiniMax
AI Summarization OpenAI GPT-4o Mini / GPT-4o
Charts Chart.js 4
PDF Export html2pdf.js
Word Export docx
Icons Tabler Icons
Audio Processing ffmpeg (fluent-ffmpeg)
Database SQLite (better-sqlite3, WAL mode)
File Upload Multer
Desktop Electron 35

Contributing

Contributions are welcome and appreciated. Here's how to get started:

  1. Fork the repository
  2. Create a feature branch: git checkout -b feature/amazing-feature
  3. Make your changes and commit: git commit -m 'feat: add amazing feature'
  4. Push to your fork: git push origin feature/amazing-feature
  5. Open a Pull Request

Please ensure your code follows the existing style conventions and includes relevant tests when applicable.


License

MIT Β© Esteban CortΓ©s


Built with OpenAI, Vue.js, Express, ffmpeg, and Chart.js

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages