Centurion

CENTURION

SPEECH IN · SUBTITLES OUT · .NET 10

Commands

Command Family

Command What it does Usage
init Interactive setup wizard — detects files, suggests the right chain init
convert Entry point: any subtitle file (SRT/VTT/ASS…) to intermediate file convert <INPUT_FILE>
asr Speech recognition: media to intermediate file asr <INPUT_FILE>
ocr Subtitle text recognition from video or images ocr <INPUT_FILE>
from-script Script timing: media + script to intermediate file from-script <INPUT_FILE> <SCRIPT_FILE>
correct Calibrate an intermediate file’s timeline & text correct <CENTURION_FILE>
translate Translate an intermediate file (LLM, glossary, target-script) translate <CENTURION_FILE> -t <LANG>
dub Media dubbing: intermediate file to dubbed WAV (Qwen3-TTS / QORA-TTS / IndexTTS) dub <CENTURION_FILE> -t <LANG>
build Exit point: intermediate file to ASS / SRT / TXT subtitles build <CENTURION_FILE>
serve REST API over the pipeline (in-process HTTP) serve [--port 8080]
validate / migrate IR schema check & upgrade validate <FILE> or migrate <FILE> --to 1.0
models Model registry management — list / install / verify / remove models list
providers Provider capabilities — list / test providers test
pipeline-graph Visualize the DAG for any workflow pipeline-graph <FILE>
quality Per-run quality report (coverage, drift, CPS, warnings) quality <FILE>
doctor Environment diagnostics — probe and write a diagnostic log doctor

convert

Turn any supported subtitle file (SRT, VTT, ASS…) into a Centurion intermediate file. This is the universal gateway: everything else in the command family reads .centurion.json.

Centurion convert <INPUT_FILE> [options]

Options

Examples

Centurion convert subtitles.srt
Centurion convert episode.ass -o episode.centurion.json

asr

The bread and butter: media file in, intermediate file out (then build renders the ASS).

Centurion asr <INPUT_FILE> [options]

The ASR pipeline runs audio conversion preprocessing optional vocal separation transcription optional diarization splitting cleaning optional forced alignment.

Options

Examples

# The classic: transcribe build ASS in two steps
Centurion asr demo.mp4 --language en
Centurion build demo.centurion.json

# Cloud ASR: OpenAI Whisper
Centurion asr demo.mp4 -t openai --asr-api-key <KEY> --language en

# Cloud ASR: Groq (fast & cheap)
Centurion asr demo.mp4 -t groq --asr-api-key <KEY> --language en

# Cloud ASR: Alibaba DashScope (great for Chinese)
Centurion asr lecture.wav -t dashscope --asr-api-key <KEY> --language zh

# Chinese lecture, custom output
Centurion asr lecture.wav -o lecture.centurion.json --language zh
Centurion build lecture.centurion.json -o lecture.ass

# Meeting + karaoke + speaker labels
Centurion asr meeting.mp4 --karaoke --num-speakers 2
Centurion build meeting.centurion.json

OCR — Subtitle Text Recognition

Extract burned-in subtitles, game dialogue, or other visible text from video frames or a single image.

Centurion ocr <INPUT_FILE> [options]

OCR inference can run in the cloud or locally:

Backend What it is Key?
zhipu (default) Zhipu AI GLM-OCR in the cloud --ocr-api-key required
ollama Local Ollama vision model (qwen2.5vl, llava…) none
llamacpp Local llama-server with a vision GGUF none

For video, ocr uses VideoSubFinderCli from PATH or downloads and caches it on supported x64 systems. If unavailable or no frames are found, it falls back to fixed-interval FFmpeg extraction. Images are passed directly.

# Cloud GLM-OCR
Centurion ocr episode.mkv --ocr-api-key <KEY> --language zh

# Local Ollama
Centurion ocr episode.mkv --ocr-backend ollama --ocr-model qwen2.5vl:7b

# Local llama-server
Centurion ocr movie.mp4 --ocr-backend llamacpp --ocr-base-url http://127.0.0.1:8080/v1

# A screenshot / subtitle image
Centurion ocr frame.png --ocr-backend ollama

# Override VideoSubFinder and tune fallback interval
Centurion ocr movie.mp4 --ocr-api-key <KEY> --ocr-interval 1.5 --ocr-videosubfinder-path "C:\\VideoSubFinder\\VideoSubFinderCli.exe"

OCR Options


correct

Subtitles that are close-but-not-quite? convert the subtitle first, then correct the intermediate file against the source audio and/or a reference script. Output stays an intermediate file — chain it into translate, dub, or straight to build.

Centurion correct <CENTURION_FILE> [options]

Common Options

Example

# Convert first, then correct against audio + script
Centurion convert subtitles.srt
Centurion correct subtitles.centurion.json --audio podcast.mp3 --script transcript.txt

# Timeline-only pass
Centurion correct subtitles.centurion.json --audio podcast.mp3 -s timeline-only

# Render the corrected result
Centurion build subtitles.corrected.centurion.json

from-script

Got a transcript/script and the matching media? This command aligns the script to the audio and produces an intermediate file — perfect for podcasts, interviews, and any “we already know what was said” scenario.

Centurion from-script <INPUT_FILE> <SCRIPT_FILE> [options]

Common Options

Example

Centurion from-script podcast.mp3 transcript.txt --language en --max-chars-per-line 20
Centurion build podcast.centurion.json

build

Render any intermediate file into a polished ASS subtitle file. This is the only command that produces subtitles — everything upstream speaks .centurion.json.

Centurion build <CENTURION_FILE> [options]

Options

What it renders

Examples

Centurion build movie.centurion.json
Centurion build movie.corrected.centurion.json -o movie_final.ass