Industry Insights 22 min read

Open-Source Video Pipeline: Concat for Editing, capcut-cli for AI Rough-Cuts, Voice-Pro for Dubbing

This article analyzes three open-source projects that together create a fully local video production pipeline: Concat, a cross-platform CapCut alternative with GPU rendering and HDR support; capcut-cli, a CLI tool that lets AI generate editable CapCut drafts for rough cuts; and Voice-Pro, a WebUI that automates video download, vocal separation, transcription, translation, and voice cloning for dubbing.

Geek Labs
Geek Labs
Geek Labs
Open-Source Video Pipeline: Concat for Editing, capcut-cli for AI Rough-Cuts, Voice-Pro for Dubbing

Concat: A Local CapCut Alternative Built in Rust

Concat (jub0t/Concat) is a cross-platform video editor (Windows, macOS, Linux, Android, iOS via sideload) written in Rust with Tauri and React. In 25 days it released 5 versions (v0.2.0 to v0.2.5 Beta), gaining 3,867 stars. It targets the same core tasks as CapCut: auto-subtitles (Whisper), text-to-speech, chroma key, keyframe animation, styled captions, multi-track editing, and 4K export. Key differentiators: no watermark, no account, no subscription, all assets stay local. Models download once in settings then run fully offline.

It also exposes JSON-RPC, gRPC, MCP, and a CLI so AI agents can drive editing programmatically. The timeline supports multiple timelines per project, sharing media across versions. Audio cleanup (denoise, vocal enhancement, loudness normalization) and voice effects (chipmunk, robot, telephone) are built in.

Export supports H.264, HEVC, AV1 up to 4K 60 fps 10-bit with hardware encoding via NVIDIA NVENC, Intel Quick Sync, and AMD AMF.

Two Major Advances in This Sprint

GPU render pipeline: Previously CPU-only; now decode, effects, and export run on GPU. Official benchmarks show export time nearly halved. Timeline scrubbing uses magnetic snap and ripple delete; frame interpolation handles dropped frames and VFR screen recordings.

HDR end-to-end: Timeline color space selectable between SDR, HLG, PQ. HDR media decoded at 10-bit, processed on GPU, exported with BT.2020 gamut and MaxCLL/MaxFALL metadata, or tone-mapped to SDR. Includes professional scopes (waveform, RGB parade, vectorscope, histogram), lift/gamma/gain color wheels, RGB curves, and .cube LUT import (DaVinci Resolve compatible).

Mobile apps are now dedicated touch-first editors; iOS runs at 120 Hz ProMotion. Updates are delivered in-app with SHA-256 verification. License changed from MPL-2.0 to AGPL-3.0, restricting closed-source commercial forks. README pledges "no paywall, ever"; revenue comes from sponsorships.

Concat subtitle and voice cloning UI
Concat subtitle and voice cloning UI

Concat's subtitle and voice cloning interface

capcut-cli: AI Generates Editable CapCut Drafts

capcut-cli (renezander030/capcut-cli, 806 stars, MIT, Node ≥18, zero deps) solves the "AI black box" problem: instead of outputting a flat video, it writes a CapCut draft JSON that opens in CapCut/CapCut Intl for full track-level tweaking. It reads/writes the local draft file directly (schema-validated, atomic write with .bak backup), no network, no background service.

AI handles deterministic chores: rough cut, subtitles, translation, long-to-short clipping. Final export decision stays human. Granular operations include:

Inspect: info/tracks/materials/lint — dump structure, tracks, materials, potential issues.

Create: init/quickstart/compile — generate a project from scratch; quickstart adds video + subtitle track in one command.

Edit: trim, speed, transitions, masks, text/image animations, easing curves — all structured timeline ops.

Subtitles: local Whisper alignment, SRT/VTT import/export, whole-draft multi-language cloning (translation via Anthropic API).

Long-to-short: scene detection, silence detection, duplicate segment detection — for podcast/video slicing.

Effects: chroma key, smart background removal, sound effects.

Templates/presets: reusable layouts, portable text styles across projects.

Media ingestion checks Wikimedia image licenses to avoid commercial clearance issues. Output drafts open in CapCut with every track intact. Preview renders low-res proxies via FFmpeg; final quality comes from CapCut export.

Quickstart:

capcut quickstart my-first --video clip.mp4 --srt captions.srt

CJK Subtitle Handling (v0.24–v0.26)

Standard subtitle specs (42 chars/line, 20 chars/sec) are designed for Latin scripts. Whisper tokenizes Chinese/Japanese into single characters or short words, producing fragmented lines if split by English rules. capcut-cli realigns by character: Chinese 16 chars/line, Japanese 13, Korean 16, with per-language CPS limits. Lint flags overflow; --fix reflows at character boundaries. v0.26 adds character-level CJK alignment and word-by-word karaoke captions.

capcut-cli demo: vertical video with word-by-word karaoke captions
capcut-cli demo: vertical video with word-by-word karaoke captions

capcut-cli demo: vertical video with word-by-word karaoke captions

Four Integration Layers

Lightest: Install as an AI coding assistant skill ( npx skills add renezander030/capcut-cli) — works with Claude Code, Codex, Cursor. Agent learns commands, draft locations, and "inspect before mutate" habit.

Library: Typed Python/Node API, zero deps. Example Python:

from capcut import run
run("quickstart", "my-short", video="clip.mp4", ratio="9:16")

Batch: capcut serve reads JSONL from stdin, integrates with n8n, Make, Coze for pipeline production.

Sandboxed (experimental): Core compiled to WASM component with only three read-only tools (inspect, diff, lint) — no FS, network, or process access. For "AI reviews draft but gets zero permissions" scenarios.

Security notes: v0.17.2 and below had stable device IDs in test fixtures; v0.17.0 and below had local command injection. Fixed in 0.18.0 and 0.17.1 respectively. CapCut 6.0+ encrypts drafts; tool warns explicitly rather than failing silently. Community-maintained, unaffiliated with ByteDance; weekly releases (latest 2026-09-28) track CapCut updates. Author René runs the YouTube channel Bronze Age Banter (history documentaries), so features stem from real batch-production pain. Sister projects: draftcat (Go-based managed AI video pipeline) and skillgate (deterministic acceptance gates for agents).

Voice-Pro: One-Click Foreign-to-Chinese Dubbing Pipeline

Voice-Pro (abus-aikorea/voice-pro, 12,958 stars, GPL-3.0, Python/Gradio, Windows + NVIDIA GPU) consolidates the traditional fragmented workflow: download → vocal separation → transcription → translation → TTS → mix. A single WebUI accepts a video link (YouTube via yt-dlp) and outputs a subtitled, dubbed version.

Pipeline stages:

Download & separation: yt-dlp + Demucs splits vocals from background music — crucial for natural dubbing (replace vocals, keep music).

Transcription & subtitles: Three Whisper variants (original, Faster-Whisper, timestamped), 90+ languages, word-level highlighting.

Translation: 100+ languages, processes SRT/ASS files directly.

Speech synthesis: Edge-TTS (100+ languages, 400+ voices); for realism, zero-shot voice cloning via F5-TTS, E2-TTS, CosyVoice (few seconds of reference audio). Recent addition: Fun-CosyVoice3 (9 languages including Korean).

Model selection notes: kokoro ranked 2nd on HuggingFace TTS Arena (lightweight, fast); F5-TTS is the community favorite for Chinese naturalness with multilingual fine-tunes.

Target users: creators localizing repurposed videos, podcasters making multilingual versions, bulk subtitle processors. Chinese is a first-class citizen: UI in Simplified Chinese, F5-TTS excels at Chinese.

Value is not any single model but the integrated, install-once pipeline. Installation: git clone → start.bat auto-installs uv, Python 3.12, dependencies (minutes), then ~10 GB models (resumable). Requires 4 GB VRAM (8 GB recommended), 20 GB disk. Maintenance: update.bat syncs via lockfile; uninstaller preserves models/outputs; translation rate-limits auto-retry with backoff; failed lines reported separately with original text retained.

Voice-Pro Chinese UI: Dubbing Studio main page
Voice-Pro Chinese UI: Dubbing Studio main page

Voice-Pro Chinese UI: Dubbing Studio main page

Four tabs:

Dubbing Studio: end-to-end chain (download, denoise, subtitles, translate, dub).

Whisper Caption: subtitle specialist, 90+ langs, word timestamps, in-video preview.

Translate: batch subtitle translation, real-time speech recognition+translation (BBC live demo shows low latency).

Speech Generation: synthesis with speed/volume/pitch controls, batch podcast production with cloned voices.

Voice-Pro real-time speech recognition and translation UI
Voice-Pro real-time speech recognition and translation UI

Voice-Pro real-time speech recognition and translation UI

Default translation/dubbing uses free endpoints (Google Translate free tier, Edge-TTS); high volume/stability can plug in Azure keys. Development paused (author focuses on WeConnect cultural-exchange app); v4.0 (July 2025) is a polished exit gift: migrated from Miniconda to uv (minutes not hours), no admin/CUDA Toolkit needed, PyTorch 2.8, RTX 50 series support, persistent error UI. Mac/Linux marked "unverified". Voice cloning ethics: own voice or authorized only; GPL-3.0 applies to derivatives.

Putting It Together

The three tools form a natural assembly line: rough cut in Concat, repetitive batch edits via capcut-cli (AI produces editable drafts), multilingual dubbing via Voice-Pro. Their division of AI/human responsibility differs: Concat exposes full editor via MCP to agents; capcut-cli lets AI manipulate draft structure but reserves export for humans; Voice-Pro runs entirely on local models, no cloud. The open-source answer to video production has evolved from "build a better app" to "redefine the human-AI task boundary."

Repositories:

https://github.com/jub0t/Concat
https://github.com/renezander030/capcut-cli
https://github.com/abus-aikorea/voice-pro
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

CLIAIopen-sourceTTSvideo-editingpipelinedubbingWhisperCapCut-alternative
Geek Labs
Written by

Geek Labs

Daily shares of interesting GitHub open-source projects. AI tools, automation gems, technical tutorials, open-source inspiration.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.