dev.club โ€” where best developers and top companies connect.

dev.club โ€” where best developers and top companies connect.Invite only

Request invite
VideoLingo Logo

Connect the World, Frame by Frame

Huanshere%2FVideoLingo | Trendshift

English๏ฝœ็ฎ€ไฝ“ไธญๆ–‡๏ฝœ็น้ซ”ไธญๆ–‡๏ฝœๆ—ฅๆœฌ่ชž๏ฝœEspaรฑol๏ฝœะ ัƒััะบะธะน๏ฝœFranรงais

๐ŸŒŸ Overview (Try VL Now!)

VideoLingo combines speech recognition, subtitle translation, segmentation and dubbing in a Streamlit interface. It produces subtitle files and optionally subtitled or dubbed videos. Translation quality depends on the source audio, language and chosen models.

Key features:

The workflow combines transcription, translation, subtitle layout and dubbing in one project.

๐ŸŽฅ Demo

Dual Subtitles


https://github.com/user-attachments/assets/a5c3d8d1-2b29-4ba9-b0d0-25896829d951

Cosy2 Voice Clone


https://github.com/user-attachments/assets/e065fe4c-3694-477f-b4d6-316917df7c0a

GPT-SoVITS with my voice


https://github.com/user-attachments/assets/47d965b2-b4ab-4a0b-9d08-b49a7bf3508c

Language Support

Input languages:

๐Ÿ‡บ๐Ÿ‡ธ English ๐Ÿคฉ | ๐Ÿ‡ท๐Ÿ‡บ Russian ๐Ÿ˜Š | ๐Ÿ‡ซ๐Ÿ‡ท French ๐Ÿคฉ | ๐Ÿ‡ฉ๐Ÿ‡ช German ๐Ÿคฉ | ๐Ÿ‡ฎ๐Ÿ‡น Italian ๐Ÿคฉ | ๐Ÿ‡ช๐Ÿ‡ธ Spanish ๐Ÿคฉ | ๐Ÿ‡ฏ๐Ÿ‡ต Japanese ๐Ÿ˜Š | ๐Ÿ‡จ๐Ÿ‡ณ Chinese ๐Ÿคฉ

Dubbing languages depend on the selected TTS method.

Installation

VideoLingo supports Windows, macOS (Apple Silicon / Intel), and Linux.

Ask your local AI agent ๐Ÿค–

If you use an AI agent that can operate your computer, send it this prompt:

Install and launch GitHub's Huanshere/VideoLingo on my computer.

Windows: one-click install ๐ŸŽ‰

  1. Download Source code (zip) from the latest Release, extract it to your Desktop or another folder, and open the folder.
  2. Double-click OneKeyStart.bat and keep the window open. On the first run, it automatically installs uv, Python 3.12, the app dependencies, and FFmpeg. An internet connection is required.
  3. After installation, VideoLingo opens automatically in your browser. Enter your API URL, key, and model in the sidebar to start using it.

Install from source (Windows, macOS, Linux)

git clone https://github.com/Huanshere/VideoLingo.git && cd VideoLingo
uv run start.py

To start it later, run uv run start.py again from the VideoLingo folder. Apple Silicon uses MLX; Intel Macs use CPU recognition. By default, Intel Mac dubbing uses the new voice without the original background sound.

Docker (optional)

For a Linux NVIDIA container deployment, install Docker, a compatible GPU driver and the NVIDIA Container Toolkit. The image uses the same Python 3.12 setup and application dependencies, with CUDA 12.8.1/cu128 by default. See Docker docs for the matched CUDA 12.6 alternative and persistence settings.

docker build -t videolingo .
docker run -d -p 8501:8501 --gpus all videolingo

HTTP API (replaces Excel batch mode)

For agents and scripts, use the local HTTP API instead of the former Excel batch mode. It shares the Streamlit pipeline and processes one operation at a time using output/. Configure config.yaml and start it from the project root. The command also installs missing dependencies on first use:

uv run start.py --api

See the HTTP API guide for input, processing, progress, downloads, retries and serial batch processing. Interactive endpoint docs: localhost:8000/docs.

LLM, ASR and TTS providers

VideoLingo supports OpenAI-Like API format and various TTS interfaces:

For detailed installation, LLM configuration, and usage instructions, please refer to the documentation: English | ไธญๆ–‡

Current Limitations

  1. Background noise and language-specific alignment models affect recognition and word timestamps. Vocal separation may help. Numbers and symbols may lack reliable word timings; inspect the resulting subtitles.

  2. LLM output must satisfy the workflow's JSON structure. For failures, inspect output/gpt_log/error.json. Existing successful response caches and completed outputs can be reused on retry; changing the model alone does not regenerate every completed stage. Do not delete all output as the first troubleshooting step.

  3. Dubbing quality and timing depend on translation, the TTS service and speech rate. Speed adjustment does not guarantee natural delivery or perfect synchronization.

  4. Local recognition uses one primary recognition/alignment language per audio segment. Mixed-language speech is not guaranteed to retain accurate text and timing in every language.

  5. The dubbing workflow does not automatically assign a separate voice to each speaker.

๐Ÿ“„ License

This project is licensed under the Apache 2.0 License. Special thanks to the following open source projects for their contributions:

Qwen3-ASR, MLX Audio, whisperX, yt-dlp, json_repair, BELLE

๐Ÿ“ฌ Contact Me

โญ Star History

Star History Chart


If you find VideoLingo helpful, please give me a โญ๏ธ!

Join libs.tech

...and unlock some superpowers

GitHub

We won't share your data with anyone else.