Connect the World, Frame by Frame
English๏ฝ็ฎไฝไธญๆ๏ฝ็น้ซไธญๆ๏ฝๆฅๆฌ่ช๏ฝEspaรฑol๏ฝะ ัััะบะธะน๏ฝFranรงais
๐ Overview (Try VL Now!)
VideoLingo combines speech recognition, subtitle translation, segmentation and dubbing in a Streamlit interface. It produces subtitle files and optionally subtitled or dubbed videos. Translation quality depends on the source audio, language and chosen models.
Key features:
-
๐ฅ YouTube video download via yt-dlp
-
Word-level speech recognition and alignment with Qwen3-ASR + Qwen3-ForcedAligner
-
๐ NLP and AI-powered subtitle segmentation
-
๐ Custom + AI-generated terminology for coherent translation
-
Direct translation with optional reflection and natural rewriting
-
Subtitle segmentation with configurable length limits
-
๐ฃ๏ธ Dubbing with GPT-SoVITS, OpenAI, Edge TTS, and more
-
๐ One-click startup and processing in Streamlit
-
๐ Multi-language support in Streamlit UI
-
๐ Detailed logging with progress resumption
-
๐ Model searchbox with API auto-fetch โ search and filter from your provider's full model list
-
โฏ๏ธ Task control โ pause, resume, or stop processing at any step
The workflow combines transcription, translation, subtitle layout and dubbing in one project.
๐ฅ Demo
Dual Subtitleshttps://github.com/user-attachments/assets/a5c3d8d1-2b29-4ba9-b0d0-25896829d951 |
Cosy2 Voice Clonehttps://github.com/user-attachments/assets/e065fe4c-3694-477f-b4d6-316917df7c0a |
GPT-SoVITS with my voicehttps://github.com/user-attachments/assets/47d965b2-b4ab-4a0b-9d08-b49a7bf3508c |
Language Support
Input languages:
๐บ๐ธ English ๐คฉ | ๐ท๐บ Russian ๐ | ๐ซ๐ท French ๐คฉ | ๐ฉ๐ช German ๐คฉ | ๐ฎ๐น Italian ๐คฉ | ๐ช๐ธ Spanish ๐คฉ | ๐ฏ๐ต Japanese ๐ | ๐จ๐ณ Chinese ๐คฉ
Dubbing languages depend on the selected TTS method.
Installation
VideoLingo supports Windows, macOS (Apple Silicon / Intel), and Linux.
Ask your local AI agent ๐ค
If you use an AI agent that can operate your computer, send it this prompt:
Install and launch GitHub's Huanshere/VideoLingo on my computer.
Windows: one-click install ๐
- Download Source code (zip) from the latest Release, extract it to your Desktop or another folder, and open the folder.
- Double-click
OneKeyStart.batand keep the window open. On the first run, it automatically installs uv, Python 3.12, the app dependencies, and FFmpeg. An internet connection is required. - After installation, VideoLingo opens automatically in your browser. Enter your API URL, key, and model in the sidebar to start using it.
Install from source (Windows, macOS, Linux)
git clone https://github.com/Huanshere/VideoLingo.git && cd VideoLingo
uv run start.py
To start it later, run uv run start.py again from the VideoLingo folder. Apple Silicon uses MLX; Intel Macs use CPU recognition. By default, Intel Mac dubbing uses the new voice without the original background sound.
Docker (optional)
For a Linux NVIDIA container deployment, install Docker, a compatible GPU driver and the NVIDIA Container Toolkit. The image uses the same Python 3.12 setup and application dependencies, with CUDA 12.8.1/cu128 by default. See Docker docs for the matched CUDA 12.6 alternative and persistence settings.
docker build -t videolingo .
docker run -d -p 8501:8501 --gpus all videolingo
HTTP API (replaces Excel batch mode)
For agents and scripts, use the local HTTP API instead of the former Excel batch mode.
It shares the Streamlit pipeline and processes one operation at a time using output/.
Configure config.yaml and start it from the project root. The command also installs missing dependencies on first use:
uv run start.py --api
See the HTTP API guide for input, processing, progress, downloads, retries and serial batch processing. Interactive endpoint docs: localhost:8000/docs.
LLM, ASR and TTS providers
VideoLingo supports OpenAI-Like API format and various TTS interfaces:
- LLM: choose an OpenAI-compatible Chat Completions provider and model that can return the structured JSON required by the workflow. OpenLux is recommended; set the API URL to
https://api.openlux.ai/v1. Prefer GPT-6 Luna with model IDgpt-6-lunafor best value, GPT-6 Sol withgpt-6-solfor better quality, or Claude Opus 5.5 withclaude-opus-5-5for best quality. OpenLux relay rates are in the install docs. Configure the API URL, key and model in the sidebar. - Speech recognition: run Qwen3-ASR + ForcedAligner locally (default), or choose ElevenLabs or MAI-Transcribe-2 in the sidebar. MAI supports an Azure Speech key or an OpenRouter key entered in the sidebar; audio is sent to the selected provider and may incur charges. WhisperX is not installed by the installer; to use it as a backend, follow WhisperX (manual install).
- TTS: OpenAI, Fish Audio, SiliconFlow Fish/CosyVoice2, GPT-SoVITS, Edge TTS, F5-TTS and a custom adapter in
core/tts_backend/custom_tts.py.
For detailed installation, LLM configuration, and usage instructions, please refer to the documentation: English | ไธญๆ
Current Limitations
-
Background noise and language-specific alignment models affect recognition and word timestamps. Vocal separation may help. Numbers and symbols may lack reliable word timings; inspect the resulting subtitles.
-
LLM output must satisfy the workflow's JSON structure. For failures, inspect
output/gpt_log/error.json. Existing successful response caches and completed outputs can be reused on retry; changing the model alone does not regenerate every completed stage. Do not delete all output as the first troubleshooting step. -
Dubbing quality and timing depend on translation, the TTS service and speech rate. Speed adjustment does not guarantee natural delivery or perfect synchronization.
-
Local recognition uses one primary recognition/alignment language per audio segment. Mixed-language speech is not guaranteed to retain accurate text and timing in every language.
-
The dubbing workflow does not automatically assign a separate voice to each speaker.
๐ License
This project is licensed under the Apache 2.0 License. Special thanks to the following open source projects for their contributions:
Qwen3-ASR, MLX Audio, whisperX, yt-dlp, json_repair, BELLE
๐ฌ Contact Me
- Submit Issues or Pull Requests on GitHub
- DM me on Twitter: @Huanshere
- Email me at: team@videolingo.io
โญ Star History
If you find VideoLingo helpful, please give me a โญ๏ธ!