目录

OpenDub

An Open-Source Platform for Multimodal Intelligent Video Dubbing

多模态智能视频配音开源平台

Make video dubbing understandable, inspectable, and reusable.

Project Film · Interactive Web App · Methods · Playable Examples · Documentation

Watch the OpenDub project introduction

Watch the OpenDub project introduction

4 min 37 s · 1920 x 1080 · Chinese / English subtitles · delivery details

OpenDub is a local-first research platform for multimodal video dubbing. It turns a difficult research task into a clear, interactive experience: explain the inputs, inspect complete methods developed by the team, listen to authorized archived examples, relate hearing to observable acoustic evidence, and prepare a rights-aware local project.

OpenDub does not splice internal modules from different papers into a new, unverified model. HPMDubbing, StyleDubber, and EmoDubber remain independent, complete methods. OpenDub makes their task assumptions, evidence, and usage boundaries visible in one place.

Team-Developed Methods

OpenDub presents the team’s original work as complete methods with distinct priorities, rather than treating them as interchangeable fragments.

Method Complete-method focus Upstream source
HPMDubbing Hierarchical visual prosody: lip motion, facial affect, and scene context guide duration, pitch, energy, and emotion. Repository · paper
StyleDubber Multi-scale style learning: visual frames, phonemes, and utterance-level context support clear pronunciation and character style. Repository
EmoDubber Emotion-controllable movie dubbing: lip-related alignment, pronunciation, speaker identity, and emotion-guided generation. Repository · paper
Speaker2Dubber From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency Learning. Repository · paper
InstructDubber Instruction-based Alignment for Zero-shot Movie Dubbing. Repository · paper
HiCoDiT Hierarchical Codec Diffusion for Video-to-Speech Generation. Repository · paper
CoSyncDiT CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing. Repository · paper

What You Can Explore

Task Stage Method Atlas
OpenDub task stage OpenDub method atlas
Start with the synchronized roles of video, text, and reference speech. Inspect face, lip, environment, phoneme, prosody, and output views. Explore each complete team-developed method through its original architecture, clickable components, and source record.
Compare Workbench Evidence and Studio
OpenDub comparison workbench OpenDub studio
Relate archived video and audio to waveform, log-mel, F0, energy, and frame contacts in one synchronized record. Trace evidence, select a complete method, record authorized inputs, and export a versioned local preparation record.

The Task

Silent Video + Text + Authorized Reference Speech
                         │
                         ▼
                 One complete dubbing method
                         │
                         ▼
         Target Dubbed Speech + Dubbed Video

Video dubbing is more than reading a sentence aloud. The video carries lip motion, facial expression, scene context, and timing; text defines the intended content; authorized reference speech supplies an identity and style condition. OpenDub exposes these signals as an interactive, time-aware task rather than a black-box audio button.

Listen To Archived Examples

The following clips are authorized, team-provided historical research examples. Select a method name to open its MP4 in GitHub’s video viewer; run the local web app to inspect the same assets with synchronized playback and acoustic features. These are not fresh OpenDub runs, common-input replay, or rankings.

Human portrait case · 3.0 s Animated character case · 1.36 s
Play human portrait example Play animated character example
Reference performance · HPMDubbing · StyleDubber · EmoDubber Reference performance · HPMDubbing · StyleDubber · EmoDubber
Case record · authorization record Case record · authorization record

Comparison Workbench

Animated cinematic scene · 1.56 s Presenter and display scene · 7.8 s
Play animated cinematic example Play presenter and display example
Reference performance · HPMDubbing · StyleDubber · EmoDubber Reference performance · HPMDubbing · StyleDubber · EmoDubber
Case record · authorization record Case record · authorization record

Archived research example — not a fresh OpenDub run or a common-input ranking.

Public Scope

Available now Evidence-gated by design
Interactive task explanation, method canvases, original-paper component views, local Studio preparation, evidence records, and authorized archived examples. Fresh model execution, numerical comparison, replay, and live generation require a verified method runtime, licensed weights, authorized inputs, and a real smoke test.

This distinction is deliberate. It prevents mechanism illustrations or historical media from being misrepresented as a new inference result. See the project overview and model admission policy.

Run Locally

The interactive experience runs entirely on your machine.

pnpm install
pnpm web:dev

Open http://127.0.0.1:5173 and visit Task, Methods, Examples, Compare, Evidence, and Studio. The Studio/API workflow is also available through the local compose stack:

docker compose up --build

For the full quality gate:

make check

Documentation

Responsible Use

Use only video, text, and reference speech that you own or are authorized to process. Do not impersonate people, misrepresent generated media, or redistribute restricted source material. OpenDub is designed for local-first workflows and keeps evidence, input authorization, and runtime admission explicit.

License and Citation

New OpenDub platform code is released under Apache-2.0. Upstream methods, model weights, datasets, and example media remain subject to their own licenses and permission records. See NOTICE and CITATION.cff.

邀请码
    Gitlink(确实开源)
  • 加入我们
  • 官网邮箱:gitlink@ccf.org.cn
  • QQ群
  • QQ群
  • 公众号
  • 公众号

版权所有:中国计算机学会技术支持:开源发展技术委员会
京ICP备13000930号-9 京公网安备 11010802047560号